Facebook and MicroStrategy Partner for New Big Data Analytics

During MicroStrategy's annual European user conference in Barcelona, Guy Bayes, head of Enterprise BI at Facebook, spoke on the topic of Big Data at Facebook. He discussed the technical challenge caused by the massive scale of the data generated by its 1.1+ billion members. To interactively analyze this data from any dimension, Facebook engaged MicroStrategy to create and test a new massively-parallel in-memory analytic technology which MicroStrategy calls "EMMA" (for Extended MPP Memory Architecture). The project has been underway for one year and the current prototype configuration is successfully analyzing just over half of the Facebook dataset and is achieving an average response time of under 5 seconds. At full configuration, the system will run in a cluster of hundreds of commodity servers containing over 10TB of data in memory. Mr. Bayes suggests that, once this massive interactive analytic technology is commercialized by MicroStrategy for general availability by other companies, it will have been thoroughly and strenuously tested against the largest datasets anywhere, and should enter the market as a very mature and high performance technology.



Talend Interview Questions

  1. What is Talend ?
  2. What is difference between ETL and ELT components of Talend ?
  3. How to deploy talend projects ?
  4. What are types of available version of Talend ?
  5. How to implement versioning for talend jobs ?
  6. What is tMap component ?
  7. What is difference between tMap and tJoin components ?
  8. Which component is used to sort that data ?
  9. How to perform aggregate operations/functions on data in talend ?
  10. What types of joins are supported by tMap component ?
  11. How to schedule a talend job ?
  12. How to runs talend job as web service ? 
  13. How to Integrate SVN with Talend ? 
  14. How to run talend jobs on Remote server ? 
  15. How to pass data from parent job to child jobs through trunjob component ?
  16. How to load context variables dynamically from file/database ?
  17. How to run talend jobs in Parallel ?
  18. What is Context variables ? 
  19. How to export a talend job ? 
  20. What is the purpose of Talend Runtime ?
  21. How to use Talend job conductor ?
  22. How to send email notifications with job execution status ?
  23. How to implement Full outer join in Talend ?
I will add more questions.

keep on following..

For Talend training ,  please contact us on
Email : sureshreddy1.9989241627@gmail.com
Phone : 09989241627 / 07757886316

How Will the Future of Big Data Impact the Way We Work and Live?

The semantic web, or web 3.0, is often quoted as the next phase of the Internet. In a previous post I discussed the impact of big data on the semantic web and I mentioned that the semantic web will enable all humans as well as all internet connected devices to communicates with each other as well as share and re-use data in different forms across different applications and organizations in real-time. The future of big data takes full advantages of the semantic web and it will have a vast impact on organisations and society.
Jason Hoffman, CTO of Joyent, predicts that the future of big data will be about the convergence of data, computing and networks. The PC was the convergence of the computing and networks, while the convergence of computing and data will enable analysis performed directly on Exabytes of raw data enabling ad hoc questions to be asked on extremely large data sets.
Artificial intelligence that will match human intelligence will allow us to ask questions and finding answers more easily by simply asking natural questions to computers. Already Japanese scientists have built a super-computer that mimic the brain cell network and reached 1% of brain capacity. To achieve this that simulated a network consisting of 1.73 billion nerve cells connected by 10.4 trillion synapses. The process took 40 minutes, to complete the simulation of 1 second of neuronal network activity in real, biological, time. In the coming years these super-computers will become the standard. At the moment, users still need to know what you want to know, but in a future with such super-computers it is all about the things that you don’t know.
The real benefits will be when organizations do not have to ask questions anymore to obtain answers, but simply find the answer to question they never could have thought of. Advanced pattern discovery and categorization of patters will enable algorithms to perform the decision making for organizations. Extensive and beautiful visualizations will become more important and help organizations understand the brontobytes of data.
Big data scientists will be in very high-demand in the coming decades, as McKinsey also predicted in 2011 already. The real winners in the big data startup field however, will be those companies that can make big data so easy to understand, implement and use that big data scientists are not necessary anymore. Large corporations will always employ big data scientists, but the much large market of Small and Medium sized Enterprises do not have the money to hire expensive big data scientists or analysts. Those big data startups that enable big data for SME’s without the need to hire big data experts will have a huge competitive advantage.
The algorithms developed by those big data startups will become ever smarter, smartphones will become better and in the future anyone will have a supercomputer in its pocket that can perform daunting computing tasks in real-time and visualize it on the small screen in your hand. And with the Internet of Things and trillions of sensors, the amount of data that needs to be processed by these devices will grow exponentially.
Big data will only becomes bigger and brontobytes will become common language in the boardroom. Fortunately, data storage will also become more widely available as well as cheaper in order to cope with the vast amount of data. Brontobytes of data will become so common in boardrooms, that eventually the term big data will disappear again and big data will become just data again.
However, before we have reached that stage, the growing amount of data that is processed by companies and governments will create a privacy concern. Those organizations that stick to the ethical guidelines will survive, other organizations that will take privacy lighthearted will disappear, as privacy will be self-regulating. The problem will be however with the governments as citizens cannot simply move away from their government. Large public debates about the effects of big data on consumer privacy will be inevitable and together we have to ensure that we do not end-up in Minority Report 2.0 or in a ‘1984-setting’.
The future of big data is still unsure, as the big data era is still unfolding, but it is clear that the changes ahead of us will transform organizations and societies. Big data is here to stay and organizations will have to adapt to the new paradigm. Organization might be able to postpone their big data strategy a little bit, but we have seen that organizations that already have implemented a big data strategy, do outperform their peers. Therefore, start developing your big data strategy, as there is no time to waste if your organization also wants to provide products and services in the upcoming big data era.


Original article

Recorded Future

Predictive analytics are becoming more important as they are the most valuable analysis within big data as they help predict what someone is likely to buy, visit, do or how someone will behave in the (near) future. It uses a variety of different data sets such as historical, transactional, social or customer profile data to identify risks and opportunities. Recorded Future is a big data startup, which was founded in 2009, that focuses solely on the art of predictive analytics.
They have developed linguistic and statistical algorithms that can extract information from temporal signals on the web. They scan tens of thousands different websites ranging from high-quality news publications, public niche sources, government websites, blogs, financial databases etc to identify references to entities, such as people, groups or locations, and events in the future. The algorithms can detect different time periods when the events will occur and deliver that information to the user, including sentiment analysis on the topic.
They claim to unlock the predictive power of the web with the world’s first temporal analytics engine. They work for Fortune 500 companies, advances financial institutions and government agencies from around the world. These organisations use Recorded Future as a Software-as-a-Service or developers can tap into the API that they have developed. This API gives access to the index for analysis of online media flow that spans blogs and Twitter to mainstream news to government filings all collected in real time from public sources around the world.
Recorded Future is headquartered in Cambridge, MA, an has offices in Göteborg, Sweden and Arlington, VA. In 2009 it was founded by Erik Wistrand, Staffan Truvé and Christopher Ahlberg. Since then it has received over $ 20 million in funding from Google Ventures, IA Ventures, In-Q-Tel, Atlas Venture and Balderton Capital.
Recorded Future takes a very interesting approach to big data and to give organisations the predictive insights that help them make better decisions. Co-founder Christopher Ahlberg was named among the World’s Top 100 Young Innovators by MIT Technology Review and received the TR100 award in 2002. He also has been granted two software patents, and has multiple patents pending.


Original article

Is Big Data Changing The Business You Are In Without You Realizing It?

Throughout my career, one of the primary ways to classify a company has been its industry. Knowing a company’s industry gave you a solid start on understanding the company. From an analytics perspective, you could make highly accurate assumptions about what data an organization would have, what problems it was trying to solve, and what types of analytic processes would be beneficial for the organization’s business.
I have noticed a trend forming where these historical industry classifications and associated assumptions are becoming less and less accurate. I believe that this trend is tied in large part to the rise of analytics and big data over the past few years.
Consider Nike
Much of the general public still thinks that Nike is a clothing manufacturer. However, recent innovations at Nike tell a very different story that many consumers have not come to recognize. Nike is now in the business of collecting, storing, and analyzing data for its customers.  It isn’t unique to Nike. The same phenomenon is being repeated in one way or another at many other companies.
ImageConsider the Nike+ product line and specifically the Nike+ FuelBand. While the FuelBand is being sold and marketed by what many perceive to be a “clothing company,” it really isn’t clothing at all. It is an electronics product. Yes, Nike is in the high tech manufacturing business. The product also has associated web and smartphone applications. Nike is also in the software business. But wait, there’s more! The primary purpose of the device is to capture data on your daily activity such as how many steps you take and if you meet your daily activity goals. That data gets uploaded and stored by Nike. Nike is in the data collection and storage business. How do users interact with their data? Through reports and charts within the associated applications. Nike is in the analytics as a service business.  And some might argue they are in the health business as a result of their analytics.
By now you should get the point. Based on the above, it is clear that Nike is no longer purely in the clothing business.  In fact, they are in several businesses that have nothing to do with clothing and that have completely different challenges. What led to Nike’s success in the past won’t apply directly in these new spaces. Whether or not you’ve realized this change, Nike certainly has. Nike has been steadily developing these new capabilities as it pushes to realize its vision, which is “To bring inspiration and innovation to every athlete in the world.”
What Industry Is Your Company Really In?
There is one critical nuance that needs to be understood. The fact is that consumers won’t choose to buy the Nike+ FuelBand because it is the most fashionable or the most comfortable. Consumers will choose the FuelBand if it offers the most complete data, the best tracking reports, and the best interactive apps. The success of the product will have nothing to do with fashion or clothing at all. Rather, it will have everything to do with data and analytics. Nike knows this and has successfully transitioned to support the requirements of the new product lines.  As the ability to capture, analyze, and distribute data continues to permeate more aspects of our lives, more and more businesses will end up entering into market spaces they would not have dreamed of a few years ago.
Are there areas of your business that are really more about capturing data and using it for analytics than the underlying product itself? If so, it is important to recognize that fact and to adapt your business accordingly. It would be a critical blunder to continue to focus on what your company used to be, or what customers still perceive it to be, rather than to focus on what will actually drive your products’ success in the future. That success may have little to do with the industry you’ve traditionally been a part of.
Nike FuelBands are selling like hotcakes and consumers don’t care that what they may still consider a “clothing company” is selling the product. What those same consumers may not have taken time to realize is that Nike isn’t just a clothing company anymore. Is your organization missing the chance to make a similar transition that challenges the notions of what industry you are in?


Original article

The Datification of Our Daily Lives









datification
Although the term is ugly, “datification” is rapidly becoming a big trend in our daily lives.
Datification is about taking a process or activity that was previously invisible and turning it into data. That data can then be tracked, monitored, and optimized, leading to new opportunities — and new challenges. It’s similar in some ways to the notion of “dark data” –  data that has been ignored up until now because of technology limitations. Only in this case, it’s more like “dark activities” that are now being pushed into the light.
New technologies have enabled lots of new ways to “datify” our normal activities.
For example, my exercise is now datified. I went for a run this morning, and my Fitbit One device recorded exactly how long I ran for, how many strides I took, and how many calories I burned in the process. For the first time, it’s very easy for me to track and monitor my exercise progress.
data application
And that’s just one small example. A lot of my daily activity is now automatically tracked. My network of friends is now datified with Facebook. My network of professional connections is datified with LinkedIn. My location is datified with Foursquare. My latest random thoughts are datified on Twitter. My music preferences are datified with Spotify.
Even reading books is now datified. While I’m reading on my Kindle device, it’s actually watching me. Amazon tracks my reading data and uses it to provide useful services. For example, it knows what page I’m on, so I can easily switch between different devices. It uses my reading speed to estimate how long it’s going to take me to finish a book. And they’ve incorporated some aspects of the wisdom of the crowd idea – for example, I can choose to see which passages other people have highlighted as the most interesting.
That data is also being collected and analyzed by Amazon to optimize book sales. For example, when I recently finished a book in a series by Ken Follet, I received an email the very next morning, giving me a special offer on the next book in the series (interestingly, at a price that was higher than the current “normal” price…)
Amazon can, and probably will, do lots of other things with this data in the future. For example, they could work with authors to help them optimize their books, showing them which passages people find hard to read, or identify at which pages readers tend to give up on the book.
Datification is also rampant in the business world. For example, most commercial vehicles now use GPS to track and optimize journeys. Even tires are becoming datified: Pirelli has embedded sensors into truck tires that constantly beam back information about the tire pressure and temperature, and this is used to calculate tire wear, helping lengthen the lifetime of the tire and optimize preventative maintenance.
We can expect to see much more datification in the future: it makes a huge amount of sense to datify our health, for example – soon, we’ll all be wearing sensors that track our temperature, pulse, blood pressure, and so on. Doctors will be increasingly able to advise treatments not to cure us, but to prevent us from getting sick in the first place.
Education is rapidly being datified with services such as the Khan Academy. This includes the notion of “flipping” education – the pupils watch lecture videos as the homework, and do the exercises in class, where a teacher can monitor what each pupil is struggling with, and intervene as necessary, in a very individual way.
In conclusion, I believe we’ve barely scratched the surface of what is possible. As the price of connected sensors plummets, we’re going to see many other activities being datified. Rick Smolan, the journalist behind the book “The Human Face of Big Data” calls this “creating a nervous system for the world as a whole”, and datification is therefore a key part of the “big data revolution.”

Original Article : http://smartdatacollective.com/timoelliott/133001/datification-our-daily-lives