# StrataConf & Hadoop World 2012: Big Data Use Cases, Challenges, and Lessons

DevFeed: [StrataConf & Hadoop World 2012: Big Data Use Cases, Challenges, and Lessons](<https://devfeed.tech/articles/strataconf-hadoop-world-2012-31971.md>)

Original publisher: [Read original article](<https://tech.finn.no2012/11/09/strataconf-hadoop-world-2012/>)

Author: mick

Published: 2012-11-09T13:23:06Z

Content type: article

Language: en

Sources: [Finn.no](<https://devfeed.tech/sources/finn-no.md>)

Topics: [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [big-data](<https://devfeed.tech/topics/big-data.md>), [data](<https://devfeed.tech/topics/data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Algorithms](<https://devfeed.tech/topics/algorithms.md>), [Unix](<https://devfeed.tech/topics/unix.md>), [Compression](<https://devfeed.tech/topics/compression.md>)

Tags: [algorithms](<https://devfeed.tech/tags/algorithms.md>), [apache](<https://devfeed.tech/tags/apache.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [compression](<https://devfeed.tech/tags/compression.md>), [data](<https://devfeed.tech/tags/data.md>), [database](<https://devfeed.tech/tags/database.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [unix](<https://devfeed.tech/tags/unix.md>)

## AI overview

A summary of StrataConf and Hadoop World 2012 covering Big Data use cases, industry perspectives, privacy concerns, and the challenges of using Apache Hadoop. It highlights recommendations including intelligent sampling, indexing, human oversight of automated algorithms, and organizational changes around operations and privacy.

## Source excerpt

A summary of this year's Strataconf & Hadoop World. A fascinating and inspiring conference with use-cases on both sides of an ethical divide - proof that the technologies coming are game-changers in both our industry and in society. Along with some intimidating use-cases i've never seen such recruitment efforts at any conference before, from multi-nationals to the CIA. The need for developers and data scientists in Big Data is burning - the market for Apache Hadoop Market is expected to reach $14 billion by 2017. Plenty of honesty towards the hype and the challenges involved too. A barcamp Big Data Controversies labelled it all as Big Noise and looked at ways through the hype. It presented balancing perspectives from a insurance company's statistician who has dealt successfully with the problem of too much data for a decade and a hadoop techie who could provide much desired answers to previously impossible questions. Highlights from this barcamp were... One should always use intelligent samples before ever committing to big data. Unix tools can be used but they are not very fault tolerant. You know when you're storing too much denormalised data when you're also getting high compression rates on it. MapReduce isn't everything as it can be replaced with indexing. If you try to throw automated algorithms at problems without any human intervention you're bound to bullshit. Ops hate hadoop and this needs to change. Respecting user privacy is important and requires a culture of honesty and common-sense within the company. But everyone needs to understand what's illegal and why. Noteworthy (10 minute) keynotes... * The End of the Data Warehouse. They are monuments to the old way of doing things: pretty packaging but failing to deliver the business value. But Hadoop too is still flawed... Also a blog available. Moneyball for New York City. How NYC council started combining datasets from different departments with surprising results. The Composite Database, a focus on using big da