
Closed
Posted
Paid on delivery
I need an end-to-end ETL pipeline that moves data from our on-premise databases into Google Cloud, transforms it, and lands it cleanly in BigQuery. The core stack must be Airflow for orchestration, Dataproc (running PySpark) for heavy transformations, and native BigQuery SQL for final modelling and reporting layers. You will design and implement: • Secure ingestion from the on-prem source into GCS staging • Airflow DAGs that trigger Dataproc jobs, handle retries, logging and alerting • PySpark transformation scripts on Dataproc, tuned for performance and cost • BigQuery SQL models that expose the refined tables • Parameterised configuration so environments can be promoted from dev to prod without code changes Acceptance criteria: – A scheduled Airflow DAG runs end-to-end without manual intervention. – Data in BigQuery matches source row counts and key metrics for at least three historical loads. – Code, requirements, and deployment instructions are delivered in a Git-ready repository. If you have prior experience marrying Airflow, Dataproc and BigQuery at scale, I’m ready to start right away.
Project ID: 40539927
11 proposals
Remote project
Active 24 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
11 freelancers are bidding on average ₹47,205 INR for this job

Hi I will build a robast end-to-end ETL pipeline, with Airflow, DataProc and Bigquery. Please provide information on 1) Source type and data structure. 2) SQL model specifications and schema. 3) Transformation requirements. I’m happy to discuss the details over chat, and I will respond promptly. Thanks for your attention Archil
₹60,000 INR in 7 days
4.4
4.4

With your GCP ETL Pipeline Development project, choosing me as the developer will give you immediate access not only to advanced backend and database skills required for this task, but also to a wealth of experience in data processing, ETL design, and GCP architecture. My 8+ years in the industry has exposed me to numerous intricate projects demanding robust data pipelines via Airflow, large-scale data transformation on Dataproc using PySpark and BigQuery modeling. The timeliness and efficiency of my work are consistently praised by my previous clients. My proficiency in both native iOS and Android development reflects my attention to detail and adaptability, which I believe is crucial in translating your requirements into a smoothly running end-to-end project. In addition to delivering code that meets acceptance criteria; I won't just settle for completing this project. I will ensure that parameterised configurations are implemented to guarantee that your environment can be transitioned from dev to prod without having to deal with unnecessary code changes. We can even evaluate how we could integrate Firebase into the mix if it is something feasible for your project. Choose me for a well-optimised GCP ETL solution that meets both your current and future business needs.
₹38,000 INR in 7 days
1.8
1.8

Three validated historical loads is a specific requirement that deserves a clear acceptance criteria definition before the pipeline starts. Row counts plus a hash sample against the on-prem source is the baseline. For anything with business-critical data you'd typically add null checks on key columns and a uniqueness assertion at the BigQuery destination. I'd pin those criteria in the repo README so sign-off is straightforward. On the pipeline itself: Airflow DAGs with parameterized GCS bucket names, BQ dataset IDs, and Dataproc cluster configs pulled from Airflow Variables, so dev/prod promotion is a variable swap rather than file edits. Each DAG triggers an ephemeral Dataproc cluster that spins up for the job and terminates on completion, so you're not paying for idle Spark infrastructure between loads. Retry logic and alerting are in the DAG definition; dead letter handling goes to a separate GCS prefix. PySpark jobs are tuned per-job, not generic templates. On-prem schemas don't always translate cleanly to BigQuery types, so the extraction layer includes explicit type coercions and a schema validation step before any data lands in BQ. Clean git repo with dev/prod configs, documented DAGs, PySpark job modules, and three historical loads passing the agreed validation criteria. 12 days, 63,750 INR, one milestone. What's the on-prem source DB? That shapes how I write the extraction layer.
₹63,750 INR in 12 days
1.3
1.3

Hello, I can help build your end-to-end ETL pipeline on Google Cloud using Airflow, Dataproc (PySpark), and BigQuery. My experience includes designing data pipelines, building ETL/ELT workflows in Python, orchestrating jobs with Airflow, developing PySpark transformations, and creating analytical models in SQL. For this project, I can deliver: * Secure ingestion from on-prem databases to GCS staging. * Airflow DAGs with scheduling, retries, logging, monitoring, and alerting. * Optimized PySpark jobs on Dataproc for scalable transformations. * BigQuery SQL models for refined and reporting layers. * Environment-based configuration for seamless promotion between dev, test, and prod. * Data validation and reconciliation checks to ensure row counts and key metrics match source systems. Deliverables: * Git-ready repository with clean project structure. * Airflow DAGs, PySpark scripts, and BigQuery SQL models. * Requirements, configuration files, and deployment documentation. * Validation and testing procedures for historical loads. I am available to start immediately and would be happy to discuss source systems, data volume, refresh schedules, and architecture requirements. Best regards, Mahmoud Hassan
₹56,250 INR in 7 days
0.0
0.0

Getting on-prem data into BigQuery smoothly requires an Airflow and Dataproc pipeline that won't fail silently or spike compute costs. I build production-grade Python and SQL data pipelines every single day. I engineered the entire backend ETL infrastructure for my startup (Synlitics) to process live, messy inputs automatically, so I know how to handle tight logging and heavy PySpark transformations. I can map out your DAGs and have a reliable dev pipeline running in under 5 days. Happy to do a quick, free audit first so you know exactly what you're getting. What database engine is your on-premise data currently sitting in?
₹39,500 INR in 7 days
0.0
0.0

Hi there, I'm a Senior Data Engineer with 7 years of experience developing, maintaining, and optimizing ETL pipelines. I have worked extensively with Dataproc (PySpark) and self-hosted Airflow to automate workflows, implement error alerting, and ensure pipeline reliability. I have delivered data to multiple destinations, including Parquet files on GCS, BigQuery, and PostgreSQL. I also have experience optimizing ETL pipelines, achieving significant performance improvements. I can set up automated data pipelines tailored to your requirements. Let's connect to discuss your project in more detail. I'm available to start immediately. Thank you. Best regards, Charles
₹37,500 INR in 14 days
0.0
0.0

Hello, this is squarely my lane - on-prem to GCS to BigQuery, with Airflow orchestrating Dataproc PySpark jobs and BigQuery SQL for the modelling layer. I deliver it as production-ready code, not a black box - Airflow DAGs with retries, logging and alerting, parameterised PySpark transforms tuned for Dataproc, secure GCS staging ingestion from your on-prem source, and BigQuery SQL models exposing the refined tables, all driven by config so environments and tables switch cleanly. I validate the transform and SQL logic against representative sample data before deployment, so you see correct output early, then we deploy and tune cost and performance on your GCP with access shared here on Freelancer. Milestone 1 is the core path end to end - one source ingested to GCS, one Dataproc PySpark transform, and one BigQuery model, with a working Airflow DAG - then we extend to the full set of sources and models. I bid at your budget floor and keep all work and communication on Freelancer, starting once the first milestone is funded. Happy to align on your source database and target schema first. Thanks, Ricardo
₹37,500 INR in 14 days
0.0
0.0

I have 4+ years of experience building scalable ETL/ELT pipelines on Google Cloud using Python, SQL, BigQuery, Airflow, and PySpark. I have hands-on experience designing end-to-end data migration and transformation workflows from on-prem systems to GCP. I can build a robust pipeline to securely ingest data into GCS, orchestrate workflows using Airflow DAGs with retries, logging, and alerts, perform large-scale transformations using PySpark on Dataproc, and model refined datasets in BigQuery. I will ensure parameterized configurations for seamless dev-to-prod promotion, complete validation against source metrics, and deliver production-ready Git-managed code with deployment documentation.
₹56,250 INR in 7 days
0.0
0.0

Hi there, As an IIT graduate with 6 years of corporate data engineering experience, this exact stack—Airflow, Dataproc (PySpark), and BigQuery—is my core expertise. I have spent my career building and tuning the exact on-premise to GCP pipelines you need. Here is how my experience directly addresses your requirements: Secure Ingestion & GCS Staging: Proven track record of safely extracting on-premise databases into Google Cloud. Dataproc & PySpark: I write highly optimized PySpark scripts for heavy transformations, complex data filtering, and reliable data accumulation. I focus heavily on performance tuning to keep your compute costs low. Airflow Orchestration: I build resilient, fully parameterized Airflow DAGs equipped with retries, logging, and alerting to ensure zero-intervention end-to-end execution. BigQuery SQL Modeling: I design native SQL models to cleanly expose refined data for your reporting layers. Seamless CI/CD: My codebases are strictly parameterized, allowing safe environment promotion (dev to prod) via Git without altering the underlying code. I will fully satisfy your acceptance criteria. You will receive a scheduled, automated Airflow DAG, validated historical data loads matching source metrics perfectly, and a clean Git-ready repository with clear deployment instructions. I am ready to start immediately and deliver a robust, production-grade pipeline.
₹38,000 INR in 8 days
0.0
0.0

Hyderabad, India
Member since Jun 25, 2026
₹600-1500 INR
€750-1500 EUR
$10-45 USD
€30-250 EUR
₹400-750 INR / hour
$8-15 CAD / hour
₹37500-75000 INR
$10-30 USD / hour
$10-30 USD
₹1500-12500 INR
$750-1500 USD
$250-750 USD
₹70000-110000 INR
₹600-1500 INR
₹75000-150000 INR
$8-15 USD / hour
£250-750 GBP
₹12500-37500 INR
$10-30 USD
$750-1500 USD