Showing posts with label Enterprise Hadoop. Show all posts
Showing posts with label Enterprise Hadoop. Show all posts

Tuesday, July 2, 2013

Hadoop Hindsight #1 Start Small


I thought we would start series on some lessons we've learned.  Many of the topics I've learned the hard way so I hope it will be helpful for those a few steps behind in the journey.  YMMV, but I wish this ideology was firmly ensconced when we started.
Identify a business problem that Hadoop is uniquely suited for.
Just because you found this cool new hammer doesn't mean everything is a nail.  Find challenges that your existing tech can't answer easily.  One of our first projects involved moving 300 gigs of EDI transaction files.  A business unit was having BA's grep for customer strings on 26,000 files to find 4 or 5 files, then FTP'ing those to their deskptop for manual parsing and review.  They might spend a few HOURS doing this for each request.  It was a natural and simple use of Hadoop.  We learned a lot about design patterns, scheduling, and data cleanup.
Solve this one business challenge well.
Notice I didn't say nail it perfectly.  There are many aspects of Big Data that will challenge the way you've looked at things the last 20 years.  The solution should be good, but not necessarily perfect.  Accepting this gives time to establish PM strategy and basic design patterns.
Put together a small team that has worked well together in the past.
This is critical to your success! Please, please, please take note!  Inter-team communication is the foundation upon which your Hadoop practice will grow.  In The Mythical Man-Month my man Fredrick Brooks said:
To avoid disaster, all the teams working on a project should remain in contact with each other in as many ways as possible...
Ideally a team should consist of the following:
1 Salesman (aka VPs)
1 Agile-trained PM
1 Architect
2 Former DBAs
1-3 skilled java developers
1 Cluster Admin
Obviously this is very simplified and some roles can overlap.  My point is you should have no more than 10 people max starting out!
Support your solution.
This very same team should also live thru at least 3 months of support of the solution they've created.  Valuable insight is gained once you have to fix a few production problems.  Let the solution mature in production a bit to understand support considerations. This gives you time to adjust your design patterns. Trust me, you'll want time to reflect on your work and correct flaws.
Smash your solution and rebuild (Optional - If time permits)
Good luck getting the time, but if you're serious about a sustainable Enterprise Hadoop solution this should be rightly considered.
Go forth and multiply.
By this time your patterns and procedures should form the DNA of your new Hadoop cell. You're team should naturally develop into the evangelists and leaders upon which the mitosis of a new project occurs, carrying with it the new replicated chromosomes.  As your project cells divide and multiply, you'll be able to take on more formidable challenges.
That's all I have to say about that.

Saturday, January 26, 2013

Why Enterprise Hadoop is different than Web Hadoop.

Bringing the power of Hadoop to the enterprise is a tricky matter.  While we all know the wonderful virtues of distributed storage and compute and how it's solving Big Data problems in the web world, it is entirely a different matter when dealing with the challenges of a large enterprise.

I’m actually somewhat envious of some company’s laser-beam approach to Big Data.  Most of these environments are challenged with volume & velocity.  Our focus leans more towards variety.  Our ETL team currently has over 5,000 unique workflows.  Many of these products are redundant, useless, and pathetic attempts of moving data by myopic projects over time.  Ineffective MDM and data modeling practices have also fed this beast over the years.  I’m not suggesting that we move all of these over to Hadoop day one, but the writing is on the wall that at some point we will run into the same mess if we’re not careful.
How does one manage this variety of unique data sources?  We’re looking at various options of Talend and custom code to allow us to manage this.  This will evolve over time, but we’re trying to look ahead.
So what do you use to manage your data ingest & emit?