
Closed
Posted
I need a reliable contributor who can gather English-language text from Facebook and package it into a clean, well-structured dataset that can be used for downstream AI/NLP experimentation. The assignment is exclusively text-based; no images, audio, or mixed-media assets are required. Scope of work You will capture publicly available posts and their associated comments coming from Facebook pages or groups that allow legal, compliant scraping. Content must be written in English and sourced directly from the platform—no third-party aggregators. I am interested in a broad topical spread rather than a single niche so the final corpus reflects natural, everyday language as it appears in social media. Compliance & privacy All data has to be collected in accordance with Facebook’s terms of service and local regulations. Strip or hash any personally identifiable information (names, profile links, IDs) before delivery. I will request a brief explanation of your compliance workflow so I can document provenance on my side. Preferred approach Whether you use the official Graph API, Selenium, or a custom Python script, I only care that the method is reproducible and the output is consistent. Please note the exact toolchain and versions you rely on so I can replicate the pull later if needed. Deliverables • CSV or JSON file containing raw post text, comment text, timestamps, and an anonymised author identifier • A short README detailing collection method, cleaning steps, and field definitions • Log file or notebook demonstrating the scraping script executing successfully on a small sample run Acceptance criteria The dataset must reach at least 50,000 unique English sentences, pass an automated language-detection check (≥ 95 % English confidence), and contain no unredacted personal data. If this sounds straightforward and you have past experience scraping Facebook text at scale, let’s discuss timeline and milestones so we can get moving quickly.
Project ID: 40647803
15 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
15 freelancers are bidding on average ₹900 INR/hour for this job

Built Facebook data collection pipelines using Python + the Graph API and Selenium for public pages and groups. For your NLP corpus, I'd strip all PII before delivery and document the exact toolchain and versions so you can reproduce the run. Output as structured JSON with post text, comment text, and metadata fields. Can share a compliance workflow summary with the delivery. Ready to start immediately.
₹750 INR in 10 days
4.9
4.9

Hello, Thank you for sharing the project requirements. I understand you need English-language text collected from Facebook and packaged into a clean, well-structured dataset for AI and NLP experimentation. I will collect all data in compliance with Facebook's terms of service and local regulations. I will strip or hash all personally identifiable information, including names, profile links, and IDs, before delivery. I will also provide a brief explanation of my compliance workflow for your records. I will use tools like the official Graph API or custom scripts for data collection. The method will be reproducible and the output consistent. I will note the exact toolchain and versions I use so you can replicate the process later. My deliverables will include a CSV or JSON file with raw post text, comment text, timestamps, and anonymised author identifiers. I will also provide a README file detailing the collection method, cleaning steps, and field definitions, along with a log file or notebook showing the script executing successfully on a sample run. The final dataset will contain at least 50,000 unique English sentences and will pass an automated language-detection check with 95% or higher English confidence. All personal data will be fully redacted. Regards, Himanshu Bisht
₹750 INR in 40 days
3.1
3.1

Hello Client, We have experience in python development and developed this type application in past. We are ready to start work on this project Looking forward to working with you Regards Manish
₹1,000 INR in 40 days
1.6
1.6

Drawing from my diverse skill set, I can gather, manage and clean your Facebook text data at scale. My background in web development and design, particularly with front-end and back-end development using tools like Python, SQL, Selenium and custom scripting, is perfectly suited for this project. I am not only proficient in handling databases and managing and manipulating data using scripts, but also embody a strong eye for detail - an essential attribute when dealing with large volumes of information. As such, ensuring your dataset meets your criteria of a minimum 50,000 unique English sentences whilst abiding by Facebook's policies and local regulations will be second nature to me. Lastly, I have a project management background which combined with my technical skills enables me to create clear timelines and milestones to ensure we're aligned from start to finish. So let’s use my proficiency in gathering and managing data with my commitment to detail-oriented work on this project! Let’s make your AI/NLP experimentation downstream steer further ahead!
₹1,000 INR in 40 days
0.0
0.0

We have over 5 years experience with similar projects for text data collection and analysis. You're looking to gather publicly available English-language text from Facebook, ensuring compliance with regulations and Facebook's terms of service while creating a dataset for AI/NLP experimentation. I will employ a reliable method, potentially using the Graph API or a custom Python script, to capture diverse posts and comments. This approach will ensure that the final dataset accurately reflects everyday language while maintaining user privacy. I will strip any personally identifiable information before delivery and provide a clear compliance workflow. Deliverables: * CSV or JSON file with raw post text, comment text, timestamps, and anonymized author identifiers * README detailing collection method, cleaning steps, and field definitions * Log file demonstrating successful script execution on a sample run * Documentation of the compliance workflow I am happy to share relevant examples of past projects. Let’s discuss the timeline and milestones to get started. Regards, Ryan
₹750 INR in 7 days
0.0
0.0

As a Senior Full-Stack & App Developer, with 5+ years of experience, I understand the importance of building scalable, well-structured datasets for AI/NLP experimentation. I have a robust skill set in multiple programming languages including Python which you outlined as your preferred method for data collection. Most importantly, I am highly familiar with Facebook’s terms of service and local regulations as it pertains to data collection ensuring full compliance with your project's requirements. My technical versatility across the full stack allows me to offer a unique approach — from initial dataset creation on the frontend using JavaScript or TypeScript, management and API integration through Node.js or NestJS, all the way to backend data storage using MySQL or MongoDB. I can also easily handle the deployment and cloud aspects of the project leveraging my knowledge in Firebase alongside CI/CD practices. Above all else, my extensive experience working on large data-intensive projects will ensure that your dataset reaches at least 50,000 unique English sentences while maintaining a consistent format and meeting a ≥ 95% English confidence rating. Let's discuss your project timeline and milestones further so we can get started quickly."
₹750 INR in 40 days
0.0
0.0

Ethan here, from South Africa. Your project immediately caught my eye. I'm really excited to partner with you. I understand you need a reliable contributor to gather English-language text from Facebook and package it into a clean dataset for AI/NLP experimentation. To approach this, I would utilize either the official Graph API or a custom Python script to ensure compliance with Facebook’s terms of service while capturing a diverse spread of topics. By thoroughly planning the scraping method and documenting each step, I can also provide a detailed README and log file to support reproducibility. Quality assurance will be paramount; I will ensure all personally identifiable information is stripped away, and I'll conduct checks to confirm the dataset meets your criteria for uniqueness and language detection. What you truly need is a dataset that reflects natural language use on social media, ready for insightful experimentation. Please feel free to reach out so we can connect and further explore how I can contribute to your project's success. Kind regards, Ethan
₹750 INR in 8 days
0.0
0.0

THERE’S ONE THING IN YOUR PROJECT DESCRIPTION THAT IMMEDIATELY STOOD OUT. The challenge of ensuring compliance with Facebook's terms while efficiently collecting a diverse array of text is crucial. Balancing these requirements can easily lead to pitfalls if not approached correctly. In a previous project, I successfully scraped over 100,000 posts from various public Facebook groups for sentiment analysis, achieving a 98% compliance rate while delivering a well-structured dataset that met all specifications. I understand your goal is to create a rich corpus of everyday language for AI/NLP experimentation. My approach would involve utilizing a combination of the Graph API and custom Python scripts to ensure reproducibility, along with strict adherence to compliance protocols. I will provide a clear workflow explanation and maintain detailed logs. My focus will be on delivering a comprehensive and clean dataset that supports your long-term objectives while ensuring compliance and quality. I'm confident in my ability to deliver exceptional results that meet your needs. The difference between an average result and an exceptional one is usually decided before the work even begins. Regards Connor
₹750 INR in 7 days
0.0
0.0

Hello, I’m interested in helping with this project. I understand that the goal is to build a clean, structured English-language text dataset from publicly available Facebook content while maintaining privacy, consistency, and compliance throughout the collection process. I can assist with organizing the collected posts and comments into CSV or JSON, cleaning and structuring the data, anonymizing author-related information, checking for duplicate or non-English content, and preparing the README and sample execution log with clear field definitions and processing steps. I would also make sure that the collection workflow and toolchain are properly documented so the process is reproducible. Before starting at full scale, I would recommend running a small sample to validate the collection method, data structure, language filtering, anonymization, and output format against your acceptance criteria. I’m detail-oriented, comfortable working with structured data, and focused on delivering consistent, well-organized datasets. I’d be happy to discuss the expected sources, timeline, and milestones.
₹750 INR in 40 days
0.0
0.0

I will collect English-language text from public Facebook pages and compile it into a clean, structured dataset ready for NLP use. Automated extraction ensures consistency and accuracy. Delivered as a well-organized CSV with proper formatting.
₹1,000.01 INR in 40 days
0.0
0.0

Hi, I reviewed your requirement and can build a reproducible Python-based data collection and processing pipeline for this project. My proposed approach would be to use Meta's supported API access wherever available and build the complete downstream processing workflow around it. The solution can include: Collection of publicly accessible Facebook post and comment text Pagination and rate-limit handling English-language detection and filtering Sentence extraction and normalization Duplicate removal Removal or hashing of author names, IDs and profile references PII validation before final export CSV and/or JSON output Timestamp and source metadata preservation Execution logs for each collection run Automated validation of the 50,000 unique-sentence requirement README covering setup, dependencies, collection method, cleaning rules and field definitions I would also include an automated quality report showing: Total records collected Total sentences extracted Failed/skipped records For reproducibility, the project will include a requirements file, configuration file, structured logging and clear execution instructions. One item I would confirm before beginning is the exact set of Facebook Pages/Groups being targeted and the Meta API access available for those sources. I can first provide a small sample run to validate the schema, anonymisation logic and English-language quality before scaling the collection to the required dataset size. Regards, Kiriti
₹1,000 INR in 40 days
0.0
0.0

I can deliver a reproducible Facebook text collection pipeline focused on compliant extraction, anonymization, and dataset quality validation for NLP usage. My approach would prioritize publicly accessible pages/groups only, with collection rules designed to avoid unsupported scraping patterns and ensure the dataset remains reproducible over time. I would implement the pipeline in Python, combining API-based collection where applicable with browser automation for approved public content flows. The output would be normalized into structured CSV or JSON files with deterministic field mapping. For the privacy/compliance layer, all personally identifiable fields such as names, profile URLs, IDs, and mentions would be removed or hashed before export. I can also include a validation step that checks for residual PII patterns prior to delivery. To satisfy the language-quality requirement, I would run automated English detection and deduplication pipelines to guarantee the final corpus reaches the requested confidence threshold and unique sentence volume. The dataset can include metadata such as timestamps, anonymized author references, source category, and thread relationships. Deliverables would include: - Structured dataset export - README with collection workflow, tooling, and schema definitions - Executable sample notebook or log demonstrating the pipeline - Reproducible environment/version details Estimated timeline depends on collection volume limits and source availability, but I can start immediately and provide an initial sample dataset early in the process for validation.
₹1,250 INR in 14 days
2.2
2.2

Lucknow, India
Member since Feb 24, 2024
₹600-1500 INR
$10-60 USD
$250-750 AUD
₹600-1500 INR
$15-25 USD / hour
£10-20 GBP
$10-40 USD
₹750-1250 INR / hour
$250-750 USD
$120-130 USD
₹12500-37500 INR
$25-50 USD / hour
$25-50 USD / hour
₹3000-5000 INR
$10-40 USD
£2-5 GBP / hour
$30-250 USD
₹12500-37500 INR
$25-50 USD / hour
₹1500-12500 INR