
Closed
Posted
Paid on delivery
I’m building a data-processing and analytics platform on AWS EMR that pulls events from Kafka, lands them in Hadoop, and serves refined datasets to downstream teams. I need an engineer who can wire everything together, tune it for performance, and leave me with maintainable jobs and clear documentation. The stack you’ll be working with revolves around: • Apache Spark for large-scale transformations • Hive for ad-hoc SQL and schema management • HBase for low-latency look-ups Your responsibilities include designing the end-to-end flow from Kafka topics into EMR, choosing optimal storage formats, writing Spark jobs, defining Hive schemas, and configuring HBase tables. I’ll rely on you for best-practice partitioning, security hardening, and job orchestration using the tools you know best. Deliverables 1. Infrastructure scripts (CloudFormation, Terraform, or similar) that spin up the EMR environment 2. Reproducible Spark jobs stored in Git 3. Hive tables with sample queries 4. Operational HBase setup and example API calls 5. A concise runbook covering deployment, monitoring, and troubleshooting When you apply, show me past work that proves you’ve handled similar Kafka → EMR → Spark/Hive/HBase pipelines at scale. Screenshots, code snippets, or short case studies are perfect; no lengthy proposal needed. I’d like to see a first working pipeline within two weeks, so let me know your availability and any questions you might have. with AWS experience is plus
Project ID: 40526215
3 proposals
Remote project
Active 9 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
3 freelancers are bidding on average ₹4,750 INR for this job

With my rich background in transforming data into meaningful business outputs using AI and cloud-based solutions, I am best suited to be your Java Kafka EMR Hadoop Data Specialist. My thorough understanding of tools like Apache Spark, Hive, and HBase will enable me to construct a reliable end-to-end flow of your data pipeline while making sure it is optimized for performance as well as being future-proof. Additionally, my proficiency in CloudFormation and Terraform will ensure efficient spin-up of your AWS EMR environment. I have a successful track record of building scalable ETL/ELT pipelines, data lakes, and data warehouses on AWS and Azure - all core skills that align perfectly with your project. Over the years, I have immensely improved operational efficiency while reducing costs by deploying intelligent systems and real-time analytics just as you seek. Not to mention, I craft machine learning models including NLP and generative AI applications that have been instrumental in providing reliable business insights. Lastly, I understand the criticality of delivering timely results while adhering to the highest quality standards. I can assure you that the deliverables will not only be completed as per schedule but will also include clear deployment instructions, robust monitoring strategies compiled in a precise runbook, ready to be implemented for troubleshooting if needed.
₹1,200 INR in 5 days
0.0
0.0

As an official enterprise under the Ministry of MSME, Government of India, I pride myself on delivering high-performance enterprise data management solutions built to rigorous global production metrics. My journey began with top-class industrial systems training at Pricol Limited and ILJIN Koregaon and with extensive experience in Java, Kafka, Hadoop, EMR and your entire stack, I’m confident that I have what it takes to masterfully complete this project for you. Drawing from my specialized technical frameworks involving optimized backend logic in Python, C++, and Rust amongst others, I am well-versed in handling large-scale data processing and analytics. Additionally, my relational database expertise including MySQL and other complex DBMS frameworks enables me to design optimal storage formats, define Hive schemas effectively and create an operational HBase setup that guarantees low-latency look-ups. Lastly, I must emphasize the value of choosing an entity-regulated freelancer like me. Operating under strict project lifecycles, I can assure precise structuring, modular code delivery and absolute adherence to deadlines. With me at your side, you can be certain that your Kafka → EMR → Spark/HBase pipelines will be perfectly implemented at scale with a strict allegiance to best industry practices. Let's get started today!
₹1,050 INR in 7 days
0.0
0.0

I have vast expertise in Hadoop, Kafka and related technologies and have been working on AWS for more than a decade. I will be able to provide complete solution with documentation as per your requirements.
₹12,000 INR in 7 days
0.2
0.2

Hyderabad, India
Payment method verified
Member since Aug 21, 2019
₹100-400 INR / hour
₹100-400 INR / hour
₹100-400 INR / hour
₹100-400 INR / hour
₹12500-37500 INR
₹37500-75000 INR
₹12500-37500 INR
$250-750 USD
₹600-1500 INR
£750-1500 GBP
$15-25 USD / hour
₹1500-12500 INR
$4000-8000 USD
$30-250 USD
₹600-1500 INR
$1500-3000 USD
₹100-400 INR / hour
$10-30 CAD
₹750-1250 INR / hour
₹400-750 INR / hour
$100-225 USD / hour
$80-100 AUD
₹600-1500 INR
₹600-1500 INR
$30-250 USD