
Closed
Posted
Arabic PDF Data Structuring & AI Search Specialist We are looking for an experienced freelancer or full-time specialist to convert one chapter from an Arabic PDF book into structured, searchable data. This is a Proof of Concept on one chapter only, not a full-book project at this stage. The task includes: Arabic text extraction. Arabic OCR cleanup. Mixed Arabic/English text handling. PDF layout analysis. Image extraction. Table extraction. Content chunking. JSON schema creation. Concept extraction. Question/exercise extraction, if available. Page-level source referencing. Preparing the data for semantic search, vector search, and RAG systems. Providing documentation and a quality report. Required experience: Previous work with Arabic PDF content. Arabic OCR. Python. PDF processing. JSON data modeling. Search-ready data preparation. Embeddings, semantic search, or RAG experience preferred. Deliverables: Structured JSON files. Extracted images and tables. Search-ready chunks. Sample queries or a simple demo. Methodology documentation. Quality report. Please apply with: Previous Arabic PDF/OCR examples. Tools you will use. Timeline. Cost. Sample JSON schema. Explanation of your approach. Important: This is only a test project for one chapter from one Arabic book. A larger project may be discussed later depending on the quality of the output.
Project ID: 40466381
14 proposals
Remote project
Active 22 secs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs