Monday, October 19, 2020

How IBM is benefited using Artificial Intelligence?

Hello Learners...

    This post is all about," How IBM is using Artificial Intelligence in its core sectors? "

But before going directly to that point I would like to give you a brief knowledge about Artificial Intelligence and its Applications.


What is artificial intelligence?

In computer science, the term artificial intelligence (AI) refers to any human-like intelligence exhibited by a computer, robot, or other machine. In popular usage, artificial intelligence refers to the ability of a computer or machine to mimic the capabilities of the human mind—learning from examples and experience, recognizing objects, understanding and responding to language, making decisions, solving problems—and combining these and other capabilities to perform functions a human might perform, such as greeting a hotel guest or driving a car.

After decades of being relegated to science fiction, today, AI is part of our everyday lives. The surge in AI development is made possible by the sudden availability of large amounts of data and the corresponding development and wide availability of computer systems that can process all that data faster and more accurately than humans can. AI is completing our words as we type them, providing driving directions when we ask, vacuuming our floors, and recommending what we should buy or binge-watch next. And it’s driving applications—such as medical image analysis—that help skilled professionals do important work faster and with greater success.


Artificial intelligence applications


As noted earlier, artificial intelligence is everywhere today, but some of it has been around for longer than you think. Here are just a few of the most common examples:

  • Speech recognition: Also called speech to text (STT), speech recognition is AI technology that recognizes spoken words and converts them to digitized text. Speech recognition is the capability that drives computer dictation software, TV voice remotes, voice-enabled text messaging and GPS, and voice-driven phone answering menus.
  • Natural language processing (NLP): NLP enables a software application, computer, or machine to understand, interpret, and generate human text. NLP is the AI behind digital assistants (such as the aforementioned Siri and Alexa), chatbots, and other text-based virtual assistance. Some NLP uses sentiment analysis to detect the mood, attitude, or other subjective qualities in language.
  • Image recognition (computer vision or machine vision): AI technology that can identify and classify objects, people, writing, and even actions within still or moving images. Typically driven by deep neural networks, image recognition is used for fingerprint ID systems, mobile check deposit apps, video and medical image analysis, self-driving cars, and much more.
  • Real-time recommendations: Retail and entertainment web sites use neural networks to recommend additional purchases or media likely to appeal to a customer based on the customer’s past activity, the past activity of other customers, and myriad other factors, including time of day and the weather. Research has found that online recommendations can increase sales anywhere from 5% to 30%.
  • Virus and spam prevention: Once driven by rule-based expert systems, today’s virus and spam detection software employs deep neural networks that can learn to detect new types of virus and spam as quickly as cybercriminals can dream them up.
  • Automated stock trading: Designed to optimize stock portfolios, AI-driven high-frequency trading platforms make thousands or even millions of trades per day without human intervention.
  • Ride-share services: Uber, Lyft, and other ride-share services use artificial intelligence to match up passengers with drivers to minimize wait times and detours, provide reliable ETAs, and even eliminate the need for surge pricing during high-traffic periods.
  • Household robots: iRobot’s Roomba vacuum uses artificial intelligence to determine the size of a room, identify and avoid obstacles, and learn the most efficient route for vacuuming a floor. Similar technology drives robotic lawn mowers and pool cleaners.
  • Autopilot technology: This has been flying commercial and military aircraft for decades. Today, autopilot uses a combination of sensors, GPS technology, image recognition, collision avoidance technology, robotics, and natural language processing to guide an aircraft safely through the skies and update the human pilots as needed. Depending on who you ask, today’s commercial pilots spend as little as three and a half minutes manually piloting a flight.

Artificial intelligence and IBM Cloud



IBM has been a leader in advancing AI-driven technologies for enterprises and has pioneered the future of machine learning systems for multiple industries. Based on decades of AI research, years of experience working with organizations of all sizes, and on learnings from over 30,000 IBM Watson engagements, IBM has developed the AI Ladder for successful artificial intelligence deployments:

  • Collect: Simplifying data collection and accessibility.
  • Analyze: Building scalable and trustworthy AI-driven systems.
  • Infuse: Integrating and optimizing systems across an entire business framework.
  • Modernize: Bringing your AI applications and systems to the cloud.

The IBM Cloud AI services start with Watson Studio for building and training AI models, preparing data and performing analysis on the data. This is available in one integrated environment. For existing data, there is Watson Knowledge Catalog to do intelligent data and analytic asset discovery, cataloging and governance, and Watson Discovery to find connections and relations.

IBM has made a point of noting that only 20% of the world’s data is searchable, and there is heavy emphasis in IBM Cloud Watson on data processing and discovery. An example of this is IBM Watson Services for Core ML, which allows enterprises to build AI-powered apps that securely connect to their data and run either on-premises, offline or in cloud. These apps utilize machine learning to adapt and improve through each user interaction.

Other data discovery apps include Data Refinery, a self-service data preparation tool for data scientists, engineers and business analysts and Deep Learning, which helps developers design and deploy deep learning models using neural networks, easily scale to hundreds of training runs.

To build AI platforms, IBM has Watson Assistant to build and deploy chat bots and virtual assistants, Watson IoT Platform to provide a cloud-hosted service for device registration, connectivity, control, rapid visualization and data storage.

IBM is also big on language recognition and translation. Watson Speech to Text (STT) converts audio and voice into written text, while Watson Text to Speech (TTS) does the opposite and converts written text into natural-sounding audio in a variety of languages and voices.

How IBM Is Using AI At Scale To Benefit The Media Industry


IBM (NYSE: IBM) recently announced 3 new products to add to its growing suite of artificial intelligence solutions for brands and publishers. The new capabilities are privacy-forward and designed to enable brands to reach consumers while considering user privacy. And IBM intends to work with industry leaders, including Xandr/AT&T, Magnite, Nielsen, MediaMath, LiveRamp, and Beeswax to help scale the use of artificial intelligence across the industry.

The IBM Watson Advertising suite of solutions utilizes artificial intelligence to help clients make informed data-based decisions. And expanding on recent additions to the suite, including Watson Advertising Accelerator, Watson Advertising Social Targeting with Influential and Watson Advertising Weather Targeting among others, the planned AI-enabled capabilities also include:

1.) Extensions for IBM Watson Advertising Accelerator: Enhanced video and OTT capabilities available in the next few months that are expected to leverage Watson Machine Learning to help enable marketers to pivot video advertising creative based on individual user reaction.
2.) IBM Watson Advertising Attribution: The Beta Solution available in the coming months that utilizes Watson Machine Learning — which allows marketers to accurately quantify the efficacy of their advertising spend while understanding intent and performance drivers.
3.) IBM Watson Advertising Predictive Audiences: Solution utilizes Watson Discovery, to help enable marketers to progress beyond ‘look-alike’ to ‘do-alike’ segments in a privacy-forward format to reach consumers that exhibit similar behaviors.

IBM Watson is able to offer increased AI capabilities for businesses across language, automation, and trust. And the planned capabilities from IBM Watson Advertising will be designed to help infuse trust and transparency into the advertising ecosystem. IBM is also negotiating definitive agreements with Xandr/AT&T and Magnite.

IBM sees AI benefits in phase-change memory


In a development that holds promise of more sophisticated programming of mobile devices, drones and robots that rely on artificial intelligence, IBM researchers say they have devised a programming approach that achieves greater accuracy and reduced energy consumption.

AI systems generally employ procedures that divide memory and processing units. This practice means time is consumed transferring data between the two waypoints. The volume of data transfer is massive enough to accrue costly energy tabs.

Nature Communications reported this week that IBM devised an approach that relies on  to execute code faster and cheaper. This is a type of random access memory containing elements that can rapidly change between amorphous and crystalline states, offering performance superior to the more commonly used Flash memory modules. It is also known as P-RAM or PCM. Some refer to it as "perfect RAM" because of its extraordinary performance capabilities.

PCM relies on chalcogenide glass, which has a unique capacity to alter its state when a current passes through. A key advantage of phase change technology, first explored by Hewlett Packard, is that the memory state does not require continuous power to remain stable. The addition of data in PCM does not require an erase cycle, typical of other types of memory storage. Also, since code may be executed directly from memory rather than being copied into RAM, PCM operates faster.

IBM recognized that the growing requirements of operations relying on  in the fields of image and speech recognition, gaming and robotics demand greater efficiencies.

"As  continues to evolve and demand greater processing power," an IBM team studying solutions posted on a company blog, "companies with large data centers will quickly realize that building more  to support an additional one million times the operations needed to run categorizations of a single image, for example, is just not economical, nor sustainable."

"Clearly, we need to take the efficiency route going forward by optimizing microchips and hardware to get such devices running on fewer watts," the report states.

IBM compared PCM to the human brain, noting that it "has no separate compartments to store and compute data, and therefore consumes significantly less energy."

One drawback with PCMs is the introduction of computational inaccuracies due to read and write conductance noise. IBM addressed that problem by introducing such noise during AI .

"Our assumption was that injecting noise comparable to the device noise during the training of DNNs would improve the robustness of the models," the IBM report states.

Their assumption was correct. Their model achieved an accuracy of 93.7 percent, which IBM researchers say is the highest accuracy rating achieved by comparable  hardware.

IBM says more work needs to be done to obtain even higher degrees of accuracy. They are pursuing studies using small-scale convolutional neural networks and generative adversarial networks, and recently reported on their progress in Frontiers in Neuroscience.

"In an era transitioning more and more towards AI-based technologies, including internet-of-things battery-powered devices and autonomous vehicles, such technologies would highly benefit from fast, low-powered, and reliably accurate DNN inference engines," the IBM report.

Why IBM is using A.I. to find jobs for people who don’t have a college degree



With the unemployment rate at a low 3.7% and the skills shortage severe, corporations need to get creative about finding talented job candidates. IBM is among the technology giants testing new methods involving artificial intelligence to overcome the labor market challenges.

AI has been applied to the job application process directly as a method to prevent human bias in hiring decisions. Now more companies are using AI assessment tools to reverse-engineer job roles and find candidates often overlooked by recruiters.

IBM introduced its SkillsBuild platform in France in May 2019 with the goal of identifying job skills and employment opportunities for members of disadvantaged communities. It will be rolled out in Germany in the coming months, followed by India, and then IBM plans to bring the platform to the U.S. in 2020, by which time it is likely that the program will have thousands of users, the company says.

The IBM initiative provides jobseekers, including those with long-term unemployment, refugees, asylum seekers and veterans, with career fit assessments, training, personalized coaching and learning needed to reenter the workforce. SkillsBuild has partnered with several NGO partners and nonprofits to form, in IBM’s words, “a new, sustainable hiring mindset,” but it is not currently used as part of the application process for IBM jobs, specifically.    

Hiring below the college degree

Within the IBM Skillsbuild platform are AI tools like MyInnerGenius, created by San Diego area-based GreatBizTools, which designs AI products to find talent in nontraditional ways. 

MyInnerGenius is helping fill entry-level and mid-level IT roles for IBM’s New Collar program in the U.S. and SkillsBuild program in Europe. Many of these roles are being filled by applicants with no prior experience in technology, many of whom don’t have college degrees.

“There are currently more than 500,000 IT job openings, and colleges are only producing around 50,000 people with degrees in IT per year,” said Denise Leaser, president of GreatBizTools. The skills shortage has made it more necessary than ever for companies like IBM to look outside of their usual applicant pool. “IBM wants to open up IT to more people, especially people who may have never thought of or considered an IT role before.”

IBM using AI to predict employee performance


IBM may have the most forward-thinking employee performance review system around. Rather than simply judge employees on what they’ve already done, the company uses its Watson AI to predict what they’re going to do in the future.

How it works: Predicting the future is right inside of Watson’s wheelhouse. In this case it isn’t determining whether you’re going to win the lottery and quit, it’s using company data to make logical projections about individual performance.

Why it’s cool: IBM has 380,000 employees worldwide. That’s a lot of performance reviews for human managers to handle, and a massive time investment if they’re going to give them the scrutiny they deserve.

Watson could, theoretically, give each individual a comprehensive evaluation based on every scrap of information available, in a fraction of the time it would take humans. And the data generated is what it uses to fuel its predictions for future performance.

According to a report from Bloomberg, IBM claims Watson predicts future employee performance with 96 percent accuracy.

What’s next: IBM wants Watson everywhere. It’s been to the Grammys and outer space, and the next stop could be your company. But at least an AI probably won’t overlook your recent training, forget about your strong sales in the first quarter, or hold the fact that you’re a Yankees fan against you.

Hope you guys enjoyed reading and gained quiet good amount of knowledge about IBM and its benefits of using AI.

Thanks a lot for reading this post.... Stay tuned for more!!!



Tuesday, September 22, 2020

How LG Electronics is Benefitted from AWS?

 

  

By using AWS for our new IoT platform, we saved 80 percent in development cost. It also enables developers to focus on writing business logic for LG service scenarios.

Kim Kunwoo Chief of the Service Development Team, LG Cloud Center, LG Electronic

In December 2017, LG Electronics (LG) announced the launch of its LG ThinQ (ThinQ) brand to categorize all its upcoming smart products and services, which feature artificial intelligence (AI) technology. The concept of ThinQ is to embed Wi-Fi chips in LG products, which allows these products to communicate with each other while learning about their user’s behavioral patterns and environments.

The Challenge

As the sale of its Wi-Fi-enabled products increased, the activation rate for LG’s smart appliances increased concurrently, which led to a burden on its servers on-premises. This growth called for a new platform that could accommodate an increasing number of smart devices connected to LG servers. As such, LG built an Internet of Things (IoT) platform for ThinQ appliances, but experienced difficulties in securing resources for server development and operations.

Why Amazon Web Services

In 2016, LG successfully migrated its IT platform for home appliances, including more than 1,000 servers, from its data center to Amazon Web Services (AWS). The business adopted Amazon Elastic Compute Cloud (Amazon EC2) and Amazon Simple Storage Service (Amazon S3).

In 2017, LG learned about AWS IoT and its serverless architecture. LG decided to use AWS IoT based on an understanding that this would reduce management time for its IoT platform. In addition, AWS was the only cloud provider whose IoT services could meet LG’s requirements to support the ThinQ platform.

The company then began using AWS IoT Core to maintain continuous connectivity between ThinQ devices and its IoT platform. LG also started using AWS Lambda (Lambda) for serverless computing that registers ThinQ devices on the cloud, stores device data, and controls device status information.

For authentication between ThinQ devices and its IoT platform’s servers, LG uses X.509 certificates. The ThinQ platform also stores data in Amazon ElastiCache and Amazon DynamoDB, which provide historical data for services. “With the previous infrastructure, we retrieved data directly from devices to get status information. Now, the Device Shadow Service feature for AWS IoT enables the system to cache the status information from the server. This feature has made it possible to deliver better service consistency,” says Kunwoo Kim, chief of the Service Development Team at the LG Cloud Center.



AWS provides multiple services at cost effective pricing with reliability and technical support.

the following image is taken from report over best public cloud service provider given by Gartner incorporate. In which AWS cloud is leading.


The Benefits:

  • Because of the managed services and AWS IoT Core, LG was able to relaunch its IoT platform with a modest number of development resources. By using AWS for our IoT platform, LG saved 80 percent in development cost.
  • The business is working with AWS solutions architects to resolve issues with its cloud infrastructure.
  • LG is also using AWS IoT and a serverless architecture for air-quality monitoring services, customer service chatbots, and energy solutions.
  • With AWS, developers can improve efficiency by building the infrastructure themselves and then consulting with our infrastructure team when development work is almost complete.
  • Hope you find this case study interesting.




Thursday, September 17, 2020

Distributed Storage Cluster and Hadoop

Introduction to Big Data :


     There is no place where Big Data does not exist! The curiosity about what is Big Data has been soaring in the past few years. Let me tell you some mind-boggling facts! Forbes reports that every minute, users watch 4.15 million YouTube videos, send 456,000 tweets on Twitter, post 46,740 photos on Instagram and there are 510,000 comments posted and 293,000 statuses updated on Facebook!

Just imagine the huge chunk of data that is produced with such activities. This constant creation of data using social media, business applications, telecom and various other domains is leading to the formation of Big Data.

In order to explain what is Big Data, I will be covering the following topics:

  • Evolution of Big Data
  • Big Data Defined
  • Characteristics of Big Data
  • Big Data Analytics
  • Industrial Applications of Big Data
  • Scope of Big Data

Evolution of Big Data

Before exploring any further, let me begin by giving some insight into why the this technology has gained so much importance.

When was the last time you guys remember using a floppy or a CD to store your data? Let me guess, had to go way back in the early 21st century right? The use of manual paper records, files, floppy and discs have now become obsolete. The reason for this is the exponential growth of data. People began storing their data in relational database systems but with the hunger for new inventions, technologies, applications with quick response time and with the introduction of the internet, even that is insufficient now. This generation of continuous and massive data can be referred to as Big Data. There are a few other factors that characterize Big Data which I will be explaining later in this blog.

Forbes reports that there are 2.5 quintillion bytes of data created each day at our current pace, but that pace is only accelerating. Internet of Things(IoT) is one such technology which plays a major role in this acceleration. 90% of all data today was generated in the last two years.

What is Big Data?

So before I explain what is Big Data, let me also tell you what it is not! The most common myth associated with it is that it is just about the size or volume of data. But actually, it’s not just about the “big” amounts of data being collected. Big Data refers to the large amounts of data which is pouring in from various data sources and has different formats. Even previously there was huge data which were being stored in databases, but because of the varied nature of this Data, the traditional relational database systems are incapable of handling this Data. Big Data is much more than a collection of datasets with different formats, it is an important asset which can be used to obtain enumerable benefits.

The three different formats of big data are:

  1. Structured: Organised data format with a fixed schema. Ex: RDBMS
  2. Semi-Structured: Partially organised data which does not have a fixed format. Ex: XML, JSON
  3. Unstructured: Unorganised data with an unknown schema. Ex: Audio, video files etc.

Characteristics of Big Data

Following are the characteristics:


Big Data Analytics

Now that I have told you what is Big Data and how it’s being generated exponentially, let me present to you a very interesting example of how Starbucks, one of the leading coffeehouse chain is making use of this Big Data.

I came across this article by Forbes which reported how Starbucks made use of this technology to analyse the preferences of their customers to enhance and personalize their experience. They analysed their member’s coffee buying habits along with their preferred drinks to what time of day they are usually ordering. So, even when people visit a “new” Starbucks location, that store’s point-of-sale system is able to identify the customer through their smartphone and give the barista their preferred order. In addition, based on ordering preferences, their app will suggest new products that the customers might be interested in trying. This my friends is what we call Big Data Analytics.

Big Data Applications



These are some of the following domains where Big Data Applications has been revolutionized:

  • Entertainment: Netflix and Amazon use it to make shows and movie recommendations to their users.
  • Insurance: Uses this technology to predict illness, accidents and price their products accordingly.
  • Driver-less Cars: Google’s driver-less cars collect about one gigabyte of data per second. These experiments require more and more data for their successful execution.
  • Education: Opting for big data powered technology as a learning tool instead of traditional lecture methods, which enhanced the learning of students as well aided the teacher to track their performance better.
  • Automobile: Rolls Royce has embraced this technology by fitting hundreds of sensors into its engines and propulsion systems, which record every tiny detail about their operation. The changes in data in real-time are reported to engineers who will decide the best course of action such as scheduling maintenance or dispatching engineering teams should the problem require it.
  • Government: A very interesting use case is in the field of politics to analyse patterns and influence election results. Cambridge Analytica Ltd. is one such organisation which completely drives on data to change audience behaviour and plays a major role in the electoral process.

Scope of Big Data

  • Numerous Job opportunities: The career opportunities pertaining to the field of Big data include, Big Data Analyst, Big Data Engineer, Big Data solution architect etc. According to IBM, 59% of all Data Science and Analytics (DSA) job demand is in Finance and Insurance, Professional Services, and IT.
  • Rising demand for Analytics Professional: An article by Forbes reveals that “IBM predicts demand for Data Scientists will soar by 28%”. By 2020, the number of jobs for all US data professionals will increase by 364,000 openings to 2,720,000 according to IBM.
  • Salary Aspects: Forbes reported that employers are willing to pay a premium of $8,736 above median bachelor’s and graduate-level salaries, with successful applicants earning a starting salary of $80,265
  • Adoption of Big Data analytics: Immense growth in the usage of big data analysis across the world.

What is Hadoop? Introduction to Big Data & Hadoop


What is Hadoop?

Hadoop is a framework that allows you to first store Big Data in a distributed environment, so that, you can process it parallely. There are basically two components in Hadoop:


The first one is HDFS for storage (Hadoop distributed File System), that allows you to store data of various formats across a cluster. The second one is YARN, for resource management in Hadoop. It allows parallel processing over the data, i.e. stored across HDFS.

HDFS

HDFS creates an abstraction, let me simplify it for you. Similar as virtualization, you can see HDFS logically as a single unit for storing Big Data, but actually you are storing your data across multiple nodes in a distributed fashion. HDFS follows master-slave architecture.


Hadoop-as-a-Solution

Let’s understand how Hadoop provided the solution to the Big Data problems that we just discussed.

The first problem is storing Big data.

HDFS provides a distributed way to store Big data. Your data is stored in blocks across the DataNodes and you can specify the size of blocks. Basically, if you have 512MB of data and you have configured HDFS such that, it will create 128 MB of data blocks. So HDFS will divide data into 4 blocks as 512/128=4 and store it across different DataNodes, it will also replicate the data blocks on different DataNodes. Now, as we are using commodity hardware, hence storing is not a challenge.

It also solves the scaling problem. It focuses on horizontal scaling instead of vertical scaling. You can always add some extra data nodes to HDFS cluster as and when required, instead of scaling up the resources of your DataNodes. Let me summarize it for you basically for storing 1 TB of data, you don’t need a 1TB system. You can instead do it on multiple 128GB systems or even less.

Next problem was storing the variety of data.

With HDFS you can store all kinds of data whether it is structured, semi-structured or unstructured. Since in HDFS, there is no pre-dumping schema validation. And it also follows write once and read many model. Due to this, you can just write the data once and you can read it multiple times for finding insights.

Third challenge was accessing & processing the data faster.

Yes, this is one of the major challenges with Big Data. In order to solve it, we move processing to data and not data to processing. What does it mean? Instead of moving data to the master node and then processing it. In MapReduce, the processing logic is sent to the various slave nodes & then data is processed parallely across different slave nodes. Then the processed results are sent to the master node where the results is merged and the response is sent back to the client.

In YARN architecture, we have ResourceManager and NodeManager. ResourceManager might or might not be configured on the same machine as NameNode. But, NodeManagers should be configured on the same machine where DataNodes are present.

YARN



YARN performs all your processing activities by allocating resources and scheduling tasks.

Major components of Hadoop include a central library system, a Hadoop HDFS file handling system, and Hadoop MapReduce, which is a batch data handling resource. In addition to these, there’s Hadoop YARN, which is described as a clustering platform that helps to manage resources and schedule tasks. The Apache software foundation, the license holder for Hadoop, describes Hadoop YARN as 'next-generation MapReduce’ or 'MapReduce 2.0.’

Where is Hadoop used? 

Hadoop is used for:

  • Search – Yahoo, Amazon, Zvents
  • Log processing – Facebook, Yahoo
  • Data Warehouse – Facebook, AOL
  • Video and Image Analysis – New York Times, Eyealike

Till now, we have seen how Hadoop has made Big Data handling possible. But there are some scenarios where Hadoop implementation is not recommended.

When not to use Hadoop?

Following are some of those scenarios :

  • Low Latency data access : Quick access to small parts of data
  • Multiple data modification : Hadoop is a better fit only if we are primarily concerned about reading data and not modifying data.
  • Lots of small files : Hadoop is suitable for scenarios, where we have few but large files.

After knowing the best suitable use-cases, let us move on and look at a case study where Hadoop has done wonders.

Hadoop-CERN Case Study

The Large Hadron Collider in Switzerland is one of the largest and most powerful machines in the world. It is equipped with around 150 million sensors, producing a petabyte of data every second, and the data is growing continuously.

CERN researches said that this data has been scaling up in terms of amount and complexity, and one of the important task is to serve these scalable requirements. So, they setup a Hadoop cluster. By using Hadoop, they limited their cost in hardware and complexity in maintenance.

They integrated Oracle & Hadoop and they got advantages of integrating. Oracle, optimized their Online Transactional System & Hadoop provided them scalable distributed data processing platform. They designed a hybrid system, and first they moved data from Oracle to Hadoop. Then, they executed query over Hadoop data from Oracle using Oracle APIs. They also used Hadoop data formats like Avro & Parquet for high performance analytics without need of changing the end-user apps connecting to Oracle.

Here’s How Facebook Manages Big Data



Facebook analytics chief Ken Rudin says that Big Data is crucial to the company’s very being. And he has very particular ideas about how it should be managed.

“Facebook could not be Facebook without Big Data technologies,” Mr. Rudin said in an interview with CIO Journal.

Facebook’s success with Big Data isn’t shared by that many companies, though. As CIO Journal has reported, studies suggest that many companies are frustrated with the technology. Most recently, a Bain & Co. survey found that only 4% of business leaders are satisfied with the results of their Big Data efforts.

“The technology is really powerful, and people sometimes assume that any information they get out of it is valuable. For those people, it doesn’t work. Others are seeing huge benefits,” Mr. Rudin says.

Here’s how Facebook approaches Big Data, a broad term that includes the hardware and software used to analyze massive and often disparate data sets, as well as the data itself.

“I would say investments in Big Data frequently don’t pay off, but to me, that has nothing to do with technology,” Mr. Rudin said. Companies that invest in Big Data often owe their frustration to two mistakes, he said.

Mistake number one: They tend to rely too much on one technology, such as Hadoop. That isn’t to say that the technology isn’t effective, but rather that expectations about its role in an organization need to be reset.

In fact, Facebook relies on a massive installation of Hadoop, a highly scalable open-source framework that uses clusters of low-cost servers to solve problems. Facebook even designs its own hardware for this purpose. Mr. Rudin says Hadoop is just one of many Big Data technologies employed at Facebook. “Hadoop is not enough,” he says.

The analytic process at Facebook begins with a 300 petabyte data analysis warehouse. To answer a specific query, data is often pulled out of the warehouse and placed into a table so that it can be studied, he said. Mr. Rudin’s team also built a search engine that indexes data in the warehouse. These are just some of many technologies that Facebook uses to manage and analyze information.

The other major mistake that people make with Big Data, Mr. Rudin says, is that they often use it to answer meaningless questions. At Facebook, a meaningful question is defined as one that leads to an answer that provides a basis for changing  behavior.  If you can’t imagine how the answer to a question would lead you to change your business practices, the question isn’t worth asking, Mr. Rudin said. For example, one might use Big Data to determine whether women share more information online than men do. But what’s the point of knowing the answer?  “What am I going to do with that? Ask men to become women?” Mr. Rudin asked. “Unless you are looking for answers that people can do something about, it is a waste of time.”

Mr. Rudin judges the success of his analytics team in terms of the business, not in terms of the quality of the information that it produces. It doesn’t matter how brilliant the analysis might be, if the business unit can’t translate it into action. “That means we didn’t explain it well enough. It’s our fault,” he said.




Hadoop WebApp Automation

  Abstract : Today is an era of Technology and with the increase of technology the amount of data it produces increases every second even no...