
Closed
Posted
Paid on delivery
I need an end-to-end pipeline that automatically fetches full-text articles from Times of India, BBC and CNN, stores the raw and normalized data on S3, then produces concise LLM-generated summaries — all without creating duplicates when the job is rerun. Architecture & flow • An AWS Lambda “fetch” function triggers per source on a schedule, each source implemented as its own small adapter so that future outlets can be plugged in without touching the core pipeline. • The canonical URL (or another deterministic key) must become article_id so every article retains a stable identity across reruns. • Raw HTML lands in s3://news/raw/--date--, then a normalizer writes JSON to s3://news/normalized/. • A second Lambda reads the normalized JSON, calls the chosen LLM for summary output, and saves a separate JSON file in s3://news/summaries/ plus a daily manifest. • Broad categories (politics, sports, business) can derive from native site sections, rules, or an LLM — whichever is easiest to keep accurate — but finer topics such as “elections” or “IPL” must be tagged strictly through rule-based logic so users can filter by their saved interests. Data contracts version: "1.0" for normalized articles – metadata (title, author, publish_time, source, article_id) – body_html, body_text – categories[], topics[] version: "1.0" for summaries – headline, short_text, bullet_points[], model_name, prompt_version Deliverables 1. Terraform or CloudFormation template setting up the two Lambda functions, S3 buckets and IAM roles. 2. Source adapters for Times of India, BBC and CNN with accompanying unit tests. 3. Normalizer and summarizer code (Python preferred, open to other runtimes). 4. Rule set for topic extraction, documented and test-covered. 5. Sample manifests and both JSON schemas. 6. README explaining local testing, deployment and how to add a new source. Acceptance criteria • Running the job twice on the same day does not create duplicate article_ids. • Summaries respect the JSON schema and remain under 100 words. • Topics are populated solely by the supplied rules. • Adding a fourth source requires only a new adapter and config entry, no core code changes. If anything is unclear, let me know — I’d like to kick this off as soon as we agree on the approach.
Project ID: 40477400
5 proposals
Remote project
Active 23 secs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
5 freelancers are bidding on average ₹1,195 INR for this job

My passion lies in using data science and artificial intelligence to build scalable, reliable solutions that streamline operations and add value. Your News Scraper & Summarizer module project perfectly aligns with my skillset and experience as a seasoned Data Scientist & AI Automation Expert. I have successfully designed and deployed similar pipelines for major clients, fetching data from various sources and applying data normalization techniques to ensure clean and consistent outputs. With a deep understanding of both AWS and Terraform, I am confident in creating the required infrastructure including Lambda functions, S3 buckets, IAM roles; satisfying your first delivery requirement. Furthermore, my proficiency in Python- the language of choice for data manipulation and analysis-, will be invaluable in developing the source adapters for Times of India, BBC, CNN as well as the normalizer and summarizer code. My natural attention to detail ensures that your specified deliverables - correct JSON schemas, topic rules extraction - would all be met with
₹1,050 INR in 7 days
3.4
3.4

Hello, I can build this end-to-end AWS news ingestion and LLM summarization pipeline with scheduled Lambda fetchers, modular source adapters for Times of India, BBC, and CNN, deterministic article_id generation from canonical URLs, raw/normalized/summary storage in S3, duplicate-safe reruns, rule-based topic tagging, JSON schema validation, and daily manifests. I have strong experience in Python, AWS, data pipelines, LLM integrations, API design, automation, testing, and scalable architecture, so I can deliver Terraform/CloudFormation setup, adapter-based source code, unit tests, normalizer/summarizer logic, documented topic rules, sample schemas, and a clear README for local testing, deployment, and adding new sources without changing core pipeline code. Relevant Certifications: AI Fundamentals and the Cloud – AWS Python for Data Science, AI & Development – IBM Databases and SQL for Data Science with Python – IBM Getting Started with Git and GitHub – IBM Python for Everybody Specialization – University of Michigan Regards, Talha Almas Data Scientist / Data Analyst / Full Stack Dev Relevant Skills: Python, AWS Lambda, S3, IAM, Terraform/CloudFormation, LLM Summarization, OpenAI API, Data Pipelines, Web Scraping, JSON Schemas, Unit Testing, Rule-Based Tagging, Modular Architecture, Deduplication
₹1,450 INR in 1 day
3.1
3.1

At Paper Perfect, we excel at creating end-to-end data processing solutions, making us the ideal choice for your News Scraper & Summarizer module project. With proficiency in AWS Lambda, S3 storage, and Python - the key elements of your architecture, we are well-equipped to handle complex data flows with ease. We understand your need for an adaptable solution, and our experience in creating modular systems ensures that plugging in new adapters is seamless as it requires no core code changes. Moreover, our expertise goes beyond just building the code. We understand data contracts and have hands-on experience in schema designs that would keep metadata organized and summaries concise as per your requirements. Our proficiency extends to rule-based logic designs too, which means we can effectively implement fine topics tagging such as “elections” or “IPL” through set rules. Lastly, we offer full transparency and document every step of our work - like sample manifests and JSON schemas - ensuring you remain informed and in control throughout the project duration. At Paper Perfect, we don't just build software; we forge long-lasting relationships through trust and superior service. Give us a chance to bring this commitment to your project!
₹1,050 INR in 7 days
1.7
1.7

Implementing an efficient news article pipeline hinges on the proper construction of the AWS Lambda architecture. The need for stable article identifiers through deterministic keys is crucial to avoid duplicates, especially during multiple runs. A modular adapter design will enable easy integration of future sources without altering the core logic, aligning perfectly with your scalability goals. Deliverables will include Terraform templates for infrastructure deployment, robust source adapters for the specified publications, and comprehensive JSON schemas to ensure data consistency. The initial deliverable will be ready within 14 days. What's your deadline, when do you need this live?
₹925 INR in 10 days
0.0
0.0

Hi, I can help you build this AWS-based news ingestion and LLM summarization pipeline end-to-end. With experience in Python, AWS Lambda, S3, Terraform, data processing, web scraping, and AI integrations, I can deliver a scalable and duplicate-safe architecture that is easy to extend with new news sources. ~How My Skills Can Help Build modular source adapters for BBC, CNN, TOI, and future publishers Implement stable article IDs to prevent duplicate processing across reruns Develop Lambda-based fetch, normalize, and summarization workflows Configure S3 storage structure, IAM roles, and Terraform deployment Create rule-based topic extraction with unit tests and documentation Generate schema-compliant summaries and daily manifests Deliver clean code, test coverage, and clear setup documentation I understand the requirement for adapter-based extensibility, duplicate prevention, rule-only topic tagging, and AWS-native deployment. Happy to discuss the approach, timeline, and estimated effort. Looking forward to hearing from you. Thanks!
₹1,500 INR in 7 days
0.0
0.0

Mumbai, India
Payment method verified
Member since Nov 29, 2014
₹600-1000 INR
₹600-1500 INR
₹600-601 INR
₹600-1500 INR
₹600-1500 INR
$10-30 USD
₹1500-12500 INR
$250-750 USD
$30-250 USD
₹12500-37500 INR
$250-750 USD
$30-250 USD
$10-30 USD
$250-750 USD
₹80000-150000 INR
$30-250 USD
$250-750 USD
$10-20 NZD / hour
$10-30 USD
$10-30 AUD
$2-8 USD / hour
₹75000-150000 INR
$10-30 AUD
$30-250 USD
$15-25 AUD / hour