
Closed
Posted
Paid on delivery
Requirement Document Project Title: Handwriting & Pen Mark Removal from Scanned PDF/Image (Preserve Printed Content) Project Overview I am looking for an experienced Computer Vision / OCR / AI developer who can develop a solution to automatically remove all handwritten content and pen-based markings from scanned PDF documents while preserving all printed content exactly as it is. The solution should work on multilingual documents, especially: • English • Chinese (Simplified & Traditional) The final output should be a clean PDF with all printed text, tables, images and layout preserved. Functional Requirements • Remove handwritten text, notes, comments, initials, signatures, numbers and paragraphs. • Remove underlines (single/double/curved/zigzag). • Remove highlights (yellow/green/blue/pink). • Remove circles, rectangles, arrows, freehand drawings and scribbles. • Remove check marks, crosses, strike-through and corrections. • Remove margin notes, pen strokes, marker strokes and any pen/pencil annotations. • Printed characters must NEVER be removed. • Preserve printed English and Chinese text even if handwriting overlaps it. • Preserve printed tables, borders, images, logos, QR codes and barcode. • Support single and multi-page PDFs (100+ pages). • Preserve page size, layout and resolution. Edge Cases • Blue handwriting over black text • Black handwriting over black text • English handwriting on Chinese document • Chinese handwriting on English document • Small handwriting • Thick marker • Overlapping annotations 1. Development preferably in Python. 2. 2 samples are in the Word document 3. Please provide past experience in doing OCR or similar handwriting projects or functions.
Project ID: 40567783
87 proposals
Remote project
Active 19 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
87 freelancers are bidding on average $431 USD for this job

Hello, I trust you're doing well. I am well experienced in machine learning algorithms, with nearly a decade of hands-on practice. My expertise lies in developing various artificial intelligence algorithms, including the one you require, using Matlab, Python, and similar tools. I hold a doctorate from Tohoku University and have a number of publications in the same subject. My portfolio, which showcases my past work, is available for your review. Your project piqued my interest, and I would be delighted to be part of it. Let's connect to discuss in detail. Warm regards. please check my portfolio link: https://www.freelancer.com/u/sajjadtaghvaeifr
$500 USD in 7 days
7.2
7.2

Hello, I have carefully reviewed the project requirements for removing handwriting and pen marks from scanned PDFs while preserving printed content, especially in English and Chinese (Simplified & Traditional). Let's chat and discuss it further. To handle your project, I will start with developing a custom Computer Vision and OCR solution using tools like OpenCV and Tesseract. My approach will involve identifying and removing various types of handwritten content, pen-based markings, and preserving printed text, tables, images, and layout. The deliverables of the project will include a clean PDF output with all printed content intact and free from any handwritten or pen-based marks. Before signing-off my bid, I would like to ask a question, i.e., how critical is it to preserve the original layout and formatting of the documents? Best Regards, Aneesa.
$250 USD in 1 day
6.2
6.2

Hi, I understand your requirement to remove complex handwriting and annotations from multilingual PDFs while keeping printed Chinese and English text intact. I can certainly build this cleaning pipeline for you. **Experience & Approach** I recently developed a custom background removal tool and a license plate detection system using OpenCV and CNNs. For your project, I will implement a **U-Net architecture with a multi-scale attention mechanism** to segment pen strokes from printed characters. This approach excels at isolating handwritten features—even when overlapping black text—by leveraging localized color-space filtering and morphological opening operations. **Proof** In my previous image-processing projects, I achieved a 95%+ precision rate in isolating foreground objects from complex, noisy backgrounds. **Next Step** Could you share the two sample PDFs so I can test my segmentation logic against your specific pen-stroke density?
$675 USD in 7 days
6.2
6.2

Hello Sir/MAM I am a skilled full stack developer. Having rich experience in Java , C++ , C , C# , Python , Eclipse , Sql , Mysql , .Net ,Oracle , Object Oriented Programming , Data Structure , Algorithms, Linux , Windows , Cloud , Azure . I have a perfect grip on “Artificial Intelligence” “Automation” , and work in “Machine Learning” Deep Learning ”. My track record as demonstrated in my 100% job completion and 5-star review rating showcases My ability to deliver exceptional results on time and with utmost quality I believe that my skill set makes me the ideal candidate for this project Please come on chat we will discuss more about this I will be waiting for your reply . Thanks and Best Regards
$251 USD in 2 days
5.6
5.6

⭐⭐⭐⭐⭐ Remove Handwriting from Scanned PDFs While Keeping Printed Content Intact ❇️ Hi My Friend, I hope you're doing well. I’ve reviewed your project requirements and see you are looking for an AI developer to remove handwritten marks from scanned PDFs. Look no further; Zohaib is here to help you! My team has successfully completed 50+ similar projects involving Computer Vision and OCR. I will create a solution that efficiently removes all handwritten content while preserving printed text and layout. ➡️ Why Me? I can easily handle your project as I have 5 years of experience in Computer Vision, OCR, and AI development. My expertise includes image processing, data extraction, and multilingual support. Besides, I have a strong grip on various technologies like Python, TensorFlow, and OpenCV, ensuring a comprehensive approach to your project. ➡️ Let's have a quick chat to discuss your project in detail and let me show you samples of my previous work. I look forward to discussing this with you in our chat. ➡️ Skills & Experience: ✅ Computer Vision ✅ Optical Character Recognition (OCR) ✅ Image Processing ✅ Python Development ✅ Machine Learning ✅ TensorFlow ✅ OpenCV ✅ Multilingual Support ✅ PDF Manipulation ✅ Data Extraction ✅ Automation ✅ Quality Assurance Waiting for your response! Best Regards, Zohaib
$350 USD in 2 days
5.1
5.1

Hi there, I can develop a Python-based computer vision solution that automatically removes handwritten annotations from scanned PDFs while preserving the original printed content and document layout. With experience in OCR, image processing, and AI-based document analysis, I can build a pipeline that detects and removes handwriting, highlights, signatures, underlines, arrows, and other pen-based markings while protecting printed English and Chinese text, tables, images, QR codes, and barcodes. The solution will combine OCR, segmentation, and image restoration techniques to handle challenging cases such as overlapping handwriting, multilingual documents, thick markers, and mixed-color annotations. It will support single and multi-page PDFs (100+ pages), maintain the original page resolution and formatting, and generate a clean output ready for further OCR or archival purposes. I will deliver well-documented Python source code, a reproducible processing pipeline, and sample results using your provided documents. I have experience working on OCR, document preprocessing, and AI-powered image enhancement projects, and I’m confident I can build a robust solution for your requirements. Regards, Ahmad
$500 USD in 7 days
4.1
4.1

Greetings, I see that you need a solution to automatically remove handwritten notes and markings from scanned multilingual PDFs while keeping the printed content intact. This is a nuanced task that requires a solid grasp of OCR and image processing techniques, especially given the multilingual aspect with both English and Chinese. My approach would involve developing a tailored algorithm that accurately identifies and removes various forms of handwriting and markings, without affecting the printed text or layout. Utilizing Python, I would leverage machine learning and image analysis to ensure precision, particularly for edge cases like overlapping annotations and different handwriting styles. I have experience in similar OCR projects, where I've successfully implemented solutions to clean up scanned documents while preserving essential printed content. I’m confident in my ability to deliver a solution that meets your needs. Best regards, Saba Ehsan
$400 USD in 3 days
4.0
4.0

Hi, I'm a computer vision developer experienced in document image processing, OCR, and handwriting detection and removal pipelines. I've worked on separating handwritten annotations from printed content in scanned documents — including multilingual layouts with English and Chinese text — using segmentation models that distinguish pen strokes from printed characters while preserving tables, images, and layout. My work includes handling the hard edge cases: overlapping annotations on printed text, color-similar handwriting, and thick marker strokes on dense content. Looking forward to it, thanks!
$500 USD in 7 days
4.0
4.0

Hello! As per your project post, you are looking to build an AI-powered document processing solution that can accurately remove handwritten annotations and pen markings from scanned PDFs/images while preserving all printed content, including multilingual text, tables, images, and document layouts. The goal is to create a reliable computer vision pipeline capable of handling complex handwriting overlaps, highlights, signatures, and annotations across English and Chinese documents. My focus will be on architecting and developing a Python-based AI solution using the most suitable computer vision and OCR technologies based on your requirements, implementing handwriting detection, annotation segmentation, OCR-assisted printed text preservation, image restoration/inpainting, multilingual document processing, and scalable PDF processing workflows with accurate output generation. I specialize in AI/ML development, computer vision, OCR systems, document intelligence, image processing, and automation solutions. My approach focuses on building accurate and production-ready pipelines while creating a strong technical foundation for future enhancements such as batch processing, confidence scoring, cloud deployment, API integration, and automated document validation. Looking forward to your positive response. Please open your chat window for more details. Best Regards Prateek
$449 USD in 7 days
3.7
3.7

Hello There! I’m Md. Toriqul Islam, and I’m excited to partner with you. I have experience in Python, computer vision, OCR, image processing, AI models, and document automation. I understand you need an AI-based solution to remove handwritten annotations and pen markings from scanned PDFs while preserving printed English/Chinese text, tables, images, and layout. I can develop a Python-based pipeline using OCR and image processing techniques with accurate validation. I am skilled in Python, OpenCV, OCR, machine learning, PDF processing, and computer vision. I’m ready to start immediately and would be happy to review the sample documents and discuss the approach. Looking forward to hearing from you. Best regards, Md. Toriqul Islam
$250 USD in 5 days
3.5
3.5

Affordable, Early Delivery. ★★★★★★★★★★★★★★I hold a Masters degree which gives me the requisite background to handle writing from various subjects. I am a highly committed person towards my work. You can rely on QualityXenter for quality and consistency in writing. We never violate copyright rules. I have vast amount of experience in this industry since I am working from 2015 as a professional writer. I provide many modifications till to get your satisfactions. I have access to enough journals to use in your research project. I always produce quality work at VERY LOW RATES so, don't worry if you have a low budget for your work, I will be very happy to make a new client like you. I am producing quality work for my clients including ARTICLE WRITING, REPORT WRITING, ESSAY WRITING, RESEARCH PAPERS, BUSINESS PLAN, TECHNICAL WRITING, MATLAB, THESIS, ACCOUNTING & FINANCE work ETC. Go through my profile link https://www.freelancer.com/u/qualityxenter
$250 USD in 1 day
3.1
3.1

I am excited about the opportunity to develop a solution for removing handwriting from scanned multilingual PDFs. Leveraging my expertise in Computer Vision and AI, I will create a custom tool to efficiently eliminate handwritten content while preserving all printed elements precisely as required. The complexity of supporting both English and Chinese texts adds to the challenge, but I am confident in tackling various edge cases you outlined, such as overlapping annotations and varied handwriting types. I have previously worked on OCR projects similar to this, focusing on extracting and preserving text while managing complex document layouts. By utilizing advanced image processing techniques, I aim to ensure the integrity of the overall document structure is maintained, offering you a streamlined final output. To better align the solution with your expectations, What specific types of handwritten annotations should be prioritized for removal in your documents?
$250 USD in 13 days
3.2
3.2

As an adaptable and detail-oriented researcher, coder, and writer, I am well-equipped to address the complex nuances in your project. Having consistently engaged with multifaceted computational tasks ranging from coding in Matlab and Python to managing intricate data mining operations using Orange DM software, I can aptly remove handwritten markings while preserving printed content in your PDFs without missing a beat. To authenticate my proficiency, I call upon my research experience across various scientific domains such as quantum mechanics, statistical physics, and relativity that demand meticulousness much alike this project's needs. Just as I prioritize precision in my experimental setups and theoretical calculations, I will put equal dedication into ensuring each characteristic of the handwriting (be it color variations or minimal underlines) is suitably eliminated without affecting the perceptible English and Chinese text. Moreover, I have considerable experience using OCR technologies in my academic work and published papers, making me well-versed with the nuances and unique edge cases associated with handwriting removal. Let's set pen to paper, virtually!
$250 USD in 7 days
3.0
3.0

Hello! We have previously developed OCR solutions, AI document processing systems, PDF automation tools, and Computer Vision applications, and we can share relevant projects with you. We have 10+ years of experience in Python, OpenCV, OCR, Computer Vision, Deep Learning, PDF processing, and AI-based document analysis. We specialize in building intelligent document-processing solutions that preserve document quality while accurately removing unwanted annotations. Key Features: • AI-Based Handwriting Removal • English & Chinese OCR Support • Annotation & Highlight Removal • Printed Text Preservation • Table, QR & Barcode Protection • Multi-Page PDF Processing • Layout & Resolution Preservation • OpenCV & Deep Learning Models • Complete Source Code • Documentation & Demo We can develop a robust Python solution that intelligently detects and removes handwritten notes, highlights, signatures, and pen markings while preserving printed English and Chinese text, tables, images, QR codes, and the original document layout. We also have experience with OCR, PDF processing, and computer vision projects and can provide relevant examples during discussion. Thank you, Invoke Tech
$500 USD in 20 days
5.1
5.1

I’d be a strong fit for this project because it requires more than basic OCR cleanup—it needs a careful computer-vision pipeline that can separate handwriting, pen marks, highlights, and corrections from printed English/Chinese content while preserving the original document layout. I reviewed your requirement document, including the multilingual support, 100+ page PDF handling, overlap cases, and preservation rules for printed text, tables, logos, QR codes, and barcodes. My approach would combine document preprocessing, printed-content protection using OCR/layout detection, annotation/mark segmentation, color and stroke analysis, and inpainting/restoration so handwritten marks are removed while printed content remains intact. For difficult cases like black handwriting over black printed text or overlapping annotations, I would design a validation layer and confidence flags rather than blindly erasing content. The deliverables would include complete source code, installation guide, documentation, demo workflow, sample before/after outputs, and a clear explanation of the method so the system can be tested and extended reliably.
$250 USD in 7 days
2.3
2.3

Hello, I am Amna, a seasoned professional with over 6 years of experience in Image Processing, OCR, and AI development. I have carefully reviewed your project requirements for removing handwritten content from scanned PDFs while preserving printed text accurately. With a track record of successfully completing over 210 projects for a diverse range of clients, I assure you of a precise and efficient solution. My expertise covers a wide range of services, including OCR, data processing, and document conversion. I am confident in my ability to deliver a clean PDF with all printed content intact, regardless of language or complexity. Let's discuss your project further in the chat to explore how I can meet your specific needs. Regards, Amna
$250 USD in 1 day
1.1
1.1

Hello, I understand you need a system to process scanned multilingual (English/Chinese) PDFs, automatically identifying and removing all forms of handwritten annotations-from notes and highlights to complex drawings. The core challenge is isolating these markings, especially when they overlap printed content, and reconstructing a clean document that perfectly preserves the original layout, text, tables, and images. Technical approach: I will build a Python-based computer vision pipeline. Each PDF page will be converted to an image, then processed by a deep learning segmentation model (likely a U-Net architecture) trained to generate pixel-level masks for printed content versus handwritten annotations. This allows for precise removal of markings while protecting all underlying content before reassembling the cleaned images into a new PDF. Core modules: 1. PDF Ingestion & Rasterization: Handles uploads and converts multi-page PDFs into processable images. 2. Annotation Segmentation Engine: The core AI model that differentiates printed vs. handwritten elements. 3. Content Reconstruction: Applies the generated masks to remove unwanted annotations, carefully handling overlapping regions. 4. PDF Output Generator: Compiles the cleaned images back into a final, high-fidelity PDF. My implementation strategy is to start with a model focused on the most common annotation types in English to validate the core architecture. We will then expand the training dataset to include Chinese documents and all specified edge cases (highlights, scribbles, overlapping marks) to achieve comprehensive, high-accuracy results. Regards, Rohit
$250 USD in 35 days
0.8
0.8

I'll tackle the handwritten removal challenge with a tailored approach that balances feature extraction and validation. The result here will depend more on the feature and validation pipeline than on just picking a model and hoping the accuracy holds. My expertise in computer vision and machine learning, as demonstrated in my Computer Vision and ML Delivery, maps well to the delivery risk in this job. I successfully retrained automation across 30+ model classes and built a GPT-4 and FastAPI document automation pipeline that cut turnaround from 3 days to 18 hours. I'll apply a similar structure to this project, focusing on pipeline reliability, preprocessing, extraction quality, latency, reproducibility, and confidence handling. To ensure the pipeline's reliability, I'll employ a runnable Python pipeline, requirements file, README/setup guide, test examples, and confidence or fallback handling. Before delivery starts, I'll clarify scope, the first milestone, and the most important technical constraint to ensure we're on the same page. Are you fixed on the OCR stack already, or should I choose the fastest reliable option for your setup?
$412 USD in 7 days
1.0
1.0

Hello! I've built a similar solution that effectively removes handwritten content from scanned documents while preserving printed text, achieving a significant accuracy rate. My approach leverages advanced OCR and image processing techniques to ensure that printed content remains intact across multiple languages, including English and Chinese. I'm curious, how do you envision handling edge cases, like overlapping handwriting on printed text? If you're open, I can share examples of my previous work and discuss how we can tailor the solution to fit your specific needs. Let’s chat and explore the best way forward!
$500 USD in 7 days
0.0
0.0

Howdy! I've built computer vision pipelines that tackle exactly this kind of problem, combining classical image processing with deep learning segmentation to separate handwritten annotations from printed content with high precision. For your multilingual scanned PDF cleaner, my approach would layer OpenCV based preprocessing (adaptive thresholding, color channel separation to isolate ink colors like blue/red highlights) with a trained segmentation model, likely a fine tuned U Net or similar architecture, to classify each pixel as printed or handwritten. For the Chinese and English text preservation layer, I'd integrate an OCR confidence pass using PaddleOCR or Tesseract to anchor printed character regions before inpainting, so even when black handwriting overlaps black printed text, the printed content is protected. The output pipeline would reconstruct clean pages at original resolution and repack them into a properly structured PDF using PyMuPDF or pdf2image with layout and page size fully preserved. Thank you. Marcos.
$537 USD in 5 days
0.0
0.0

Ma On Shan, Hong Kong
Payment method verified
Member since Aug 23, 2015
$15-25 USD / hour
$240-2000 HKD
$8-15 USD / hour
$24000-40000 HKD
$2000-6000 HKD
₹750-1250 INR / hour
₹1500-12500 INR
$1500-3000 AUD
₹100-400 INR / hour
₹600-1500 INR
$30-40 USD / hour
₹750-1250 INR / hour
$15-25 USD / hour
₹600-1500 INR
₹750-1250 INR / hour
$250-750 USD
₹750-1250 INR / hour
£10-70 GBP
$2-8 USD / hour
₹12500-37500 INR
$250-750 USD
₹750-1250 INR / hour
₹750-1250 INR / hour
$10-30 USD
₹12500-37500 INR