Getting your Trinity Audio player ready...

In recent years, you’ve probably heard a lot about “big data” or Apache Hadoop but little of it has been enlightening or inspiring (maybe a bit mysterious like a good twist to a movie). Big data is really just a buzz word for now, but it’s what we use when we’re talking about a collection of large and complex data sets that are analyzed to reveal patterns. Some say it’s solving big problems for big businesses like Google, US Bank, and other large enterprises out there. Some say it’s too complicated for “an ordinary developer” to do. We’ll try to answer some of these questions, dispel some of these myths, and most importantly, show you how you and your business can use big data to solve your problems… big or small.

I’m not Google, why would I use it?

Big data isn’t always about solving “big” problems as much as it is about discovery. Yes, there are many case studies out there dealing with large amounts of data that benefit from technologies like Hadoop, a technology that enables you to manage and analyze data in ways you couldn’t in the past. But there are also those that could still benefit from similar toolsets, that don’t require a large dataset.

Doesn’t it cost a lot though?

Another misconception is that big data costs big dollars. Especially with advanced platforms like Amazon Web Services (AWS) and new emerging technologies like Apache Spark, this couldn’t be further from the truth. Amazon Web Services is known as a large player in site hosting, but it’s not as well known for its capabilities for processing data. Spark is an attempt to generalize Hadoop’s specific way of approaching data analysis, and advertises major speed improvements. Pairing the two enables your business to solve a world of problems with concise resources.

Yeah right…

As an arbitrary example, let’s say you support a legal entity with tons of paperwork coming through daily. All of this paperwork is scanned and turned into text using optical character recognition (OCR), but you are operating on a shoestring budget. The analysis you are running is simple, you’re building an index of words found, but you have lots of text from the massive amounts of paperwork to churn through. You likely don’t need heavy hardware to run the analysis, but using more hardware could reduce the processing time. In this case, if you ran a Hadoop cluster on AWS with 50 instances of a very basic virtual machine for two hours every night, it would cost about $100 per month. That’s not bad for 50 machines, right? Compare that to spending tens to hundreds of thousands of dollars on hardware, cooling, employees, fixing, replacing, upgrading… you get the point. Take it a step further and consider the massive amount of flexibility you have to build out the solution that makes sense for your business.

Okay, now what?

If you’re looking for a new way to bring more value to your business, consider data science and big data technologies. Replace or supplement that clunky data warehouse with an inexpensive, more flexible alternative like Hadoop. Consider the opportunity to gather more data and stop worrying about gathering too much data. (Get rid of those retention policies and focus on your archiving strategies.) In other words, try turning what you used to see as a trash can into real money for your business.