Thursday, May 15, 2014

Hadoop: Questions I Am Asking

I have close to 14 years’ experience with SQL Server for ETL around Data Warehousing. I lead a team of very talented Data Warehouse Developers who have developed and maintain a multi-terabyte data warehouse. We ETL and dimensional model data daily describing tens of thousands orders, millions of dollars of sales, millions of web site visitor metrics and tens of millions of web page views. And we do this each night, and more, and have it all ready for the C-suite execs to drink in with their morning coffee! I’m not saying this to brag (well, maybe a little), but because despite that experience, Hadoop puts me in an alien world where the normal tools of my trade don’t seem to make sense.

At this point in time, the questions I am asking myself are:

  • How much of my Data Warehouse environment and processes will eventually be replaced by Hadoop related technologies and processes?
  • What ETL processes are best done in Hadoop and which in SQL/SSIS?
  • How much of my storage will transfer to Hadoop, Archive, Raw Staged, Operational Stores and Modeled Data?
  • How big of Hadoop environment do I need to surpass the power of my current SQL environment?
  • Does Hadoop mean adapting new technology to the existing BI strategy or do we need a new BI strategy?

I am tenacious, so it not a matter of “if” but “when” I’ll know which of my old tools will work, how to use new tools and new strategies to conquer the next generation of data challenges.

18 comments:


  1. I like your writing style, it was very clear to understanding the concept well; I hope you ll keep your blog as updated.
    Regards,
    sas training in Chennai|sas course in Chennai

    ReplyDelete
  2. Thanks for sharing this unique and informative content which provided me the required information.
    Java Training in Chennai | JAVA Course in Chennai

    ReplyDelete
  3. Hi there! This article couldn't be written any higher! Reading through this put up jogs my memory of my previous roommate! He usually saved preaching approximately this. site I'll forward this post to him. Pretty sure he'll have a excellent examine. Thank you for sharing!

    ReplyDelete
  4. Crack Free Download. Notezilla lets you quickly take notes on PostIt-Esq desktop sticky notes and place them on websites, documents, folders.Notezilla Alternative

    ReplyDelete
  5. Adobe Master Collection CC 2022 Crack is a key feature of Adobe Systems that allows you to select locations instead of taking photos,.Adobe Suite (2022)

    ReplyDelete
  6. May your Christmas sparkle with moments of love, laughter and goodwill. And may the year ahead be full of contentment and joy. Have a Merry .Merry Christmas Wishes Text

    ReplyDelete
  7. The statement highlights the confusion experienced by traditional data warehouse professionals when moving to Apache Hadoop. With years of experience using Microsoft SQL Server for ETL, dimensional modeling, and managing multi-terabyte warehouses, the user is comfortable with structured relational systems where data is processed using SQL-based tools. In such environments, ETL pipelines transform and load business data—such as orders, sales, and web analytics—into well-organized schemas, enabling fast reporting and decision-making for executives.

    ReplyDelete
  8. However, Hadoop introduces a completely different approach to handling data. Instead of relying on centralized relational databases, Hadoop uses distributed storage (Hadoop Distributed File System / HDFS) and parallel processing frameworks such as Apache MapReduce and Apache Spark. This means familiar concepts like normalized tables, indexes, and traditional ETL workflows are replaced by cluster-based processing of massive structured and unstructured datasets.Big Data Projects. The challenge lies in adapting to new tools, programming models, and ways of thinking, but Hadoop provides major advantages in scalability, fault tolerance, and big data analytics that traditional systems may struggle to handle efficiently.

    ReplyDelete