
Cancelled
Posted
Software Engineering Task Author (Contract) About the Role We're looking for an experienced software engineer to design and build Frontier-style software engineering tasks — long-horizon, repository-based coding challenges used to train and evaluate AI coding agents. You'll take real engineering problems from your own area of expertise and turn them into rigorous, verifiable tasks: a seeded repository, a clear public specification, and an automated grader that checks observable behavior (builds, hidden test suites, protocol conformance, concurrency/recovery scenarios, deterministic replay, numeric budgets, etc.). This is deep, research-grade work — closer to writing a hard systems/algorithms exam with an autograder than to typical QA or content work. What You'll Do Design non-trivial engineering problems grounded in real-world systems (e.g., concurrency bugs, recovery/consistency issues, legacy modernization, protocol implementation, performance-constrained rewrites). Build a starter repository, write a complete public specification ([login to view URL]), and implement a reference solution that solves the task correctly. Write automated grading code that measures exactly what the instructions ask for — no hidden requirements, no LLM judges, no shortcuts. Produce a no-op/naive baseline to confirm the task has real difficulty, and adversarial "attack" submissions to confirm the grader can't be gamed. Validate, iterate, and submit tasks via pull request through a structured review pipeline. Use AI tools to accelerate parts of your workflow (explaining code, debugging validation errors, reviewing your own logic) — but the engineering judgment and task design must be your own. What We're Looking For Strong professional software engineering background (backend systems, distributed systems, compilers, low-level/systems programming, or similar) — ideally in an area where you have real depth and war stories, not just familiarity. Comfort with concurrency, correctness, and testing — the ability to reason precisely about what a grader should and shouldn't check. Experience writing clear technical specifications and/or authoring hard technical interview problems, coding challenges, or test suites. Git/GitHub workflow fluency (branches, PRs, CI). Self-directed: comfortable working from written guidelines, validating your own work, and iterating based on automated feedback and review. Bonus: experience with reinforcement learning, model evaluation, or "reward hacking" adversarial thinking (i.e., trying to break your own grader before someone else does). Nice to Have A specific technical niche you know cold (e.g., database internals, network protocols, embedded systems, compilers/language runtimes, numerical computing) that would make for a compelling, hard-to-fake task. Prior experience creating technical assessments, benchmarks, or eval datasets. Engagement Details Contract / freelance, task-based. Deliverables are reviewed against defined structure, CI, and quality checks before acceptance. Ongoing support and clarification provided via a shared Discord community. To apply, please share examples of complex systems you've built or debugged, any prior experience authoring technical assessments or benchmarks, and a brief note on what kind of engineering problem you'd want to turn into a task.
Project ID: 40671911
22 proposals
Remote project
Active 7 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
22 freelancers are bidding on average $9 USD/hour for this job

Hi, I specialize in turning complex engineering problems into reproducible, testable systems, so this Frontier style task creation is a strong fit. I understand the requirement is not simply to write coding exercises, but to build a seeded repository, precise specification, reference solution, robust grader, baseline, and adversarial submissions. My approach would focus on: • Designing realistic concurrency, backend, API, or distributed systems challenges • Building deterministic hidden tests with measurable acceptance criteria • Stress testing graders against naive and reward hacking solutions • Delivering everything through Git branches, PRs, and CI validation I am comfortable working with Python, Node.js, React, APIs, databases, Docker, Git, automated testing, and system architecture. Please ping me to get started and get outstanding results. Thanks!!!
$10 USD in 40 days
4.5
4.5

I design rigorous, deterministic SWE-bench style benchmark tasks featuring seeded repositories, strict behavioral autograders, and verified reference solutions. Here are the details you requested for task authoring: 1. Systems & Architecture Experience: Extensive background building and debugging high-concurrency Node.js/Python microservices, custom protocol state machines, and real-time WebSocket infrastructures. 2. Benchmark & Autograder Design: Deep familiarity with designing deterministic CI test suites (pytest/jest), edge-case validation, holdout test suites, and adversarial testing to prevent LLM reward hacking. 3. Proposed Initial Task Idea: - Topic: "Distributed In-Memory Key-Value Store with TTL & Thread-Safe Eviction" (or "WebSocket Resilient State Reconnection with Exponential Backoff"). - Challenge: Fix race conditions under concurrent worker loads and enforce deterministic memory cleanup budgets. - Grader: Multi-threaded stress tests verifying exact lock behavior, recovery from simulated network partitions, and 0% memory leaks. I can deliver the starter repository, markdown specification, reference PR, and anti-gaming test suite via your GitHub review pipeline. Let’s connect in chat to discuss your repository templates and Discord onboarding!
$8 USD in 20 days
3.7
3.7

Hello There! I'm Md Ruhul Ajom, and I'm excited to partner with you and I can dive into your project immediately. I have rich experience designing rigorous technical challenges with automated graders, reference solutions, and adversarial test cases to validate task difficulty. I understand you want a software engineering task author who can design real world grounded coding challenges, build seeded repositories with clear specifications, write automated graders that check exact observable behavior, and validate tasks with baseline and adversarial submissions through a PR review pipeline. I am skilled in systems programming, automated grading design, and technical specification writing. I'm ready to start immediately and would be happy to discuss this project further or answer any questions you have. Looking forward to hearing from you. Best regards, Ruhul Ajom
$5 USD in 40 days
4.8
4.8

With my extensive background in software engineering specializing in AI and backend development, I believe I fit the bill for your Frontier-Style Task Creator project beautifully. Over the years, I have built up my competencies in designing and implementing complex systems - a skill which is integral to creating rigorous tasks that showcase real-world engineering problems. It's not just about spotting bugs; it's about understanding the underlying issues they cause, be it recovery/consistency problems or rewriting performance-constrained codes. My deep technical knowledge in areas like database internals and embedded systems adds a unique value proposition to my task creation approach. This specific expertise enables me to bring forth exclusive problem-solving perspectives and create compelling assignments which might be hard-to-fake otherwise. Lastly, I take immense pride and ownership in my work. It's not just about writing a line of code; it's about delivering a robust solution that actually runs in production. My clients appreciate my commitment to understanding their needs thoroughly from project inception to completion. From producing clear technical specifications to ensuring consistent CI/CD workflow on Git/GitHub, quality and precision are assured. By hiring me for this critical role, you can expect industry-grade outcomes filled with deep explorations, rigorous validations, and effective automation solutions.
$2 USD in 40 days
3.0
3.0

Hi-Abror Here From Uzbekistan. "Rigorous Frontier-Style Engineering Tasks" - I can create repository-based challenges with deterministic graders, reference solutions, adversarial tests, and precise acceptance criteria. I can design difficult systems problems, prepare seeded repositories and instruction files, implement reference solutions, and build automated graders validating concurrency, recovery, protocols, performance, and correctness. I will create naive baselines, attack submissions, CI workflows, reproducible tests, and pull requests while iterating against review feedback and automated validation. What engineering domain would you like the first task to focus on? Looking forward to working with you.
$20 USD in 40 days
2.0
2.0

Hey! We are a team of 62 software engineering professionals with 9+ years of experience building complex backend systems, concurrency-heavy applications, automated test suites, and production-grade development workflows. We can design rigorous engineering tasks that are challenging, reproducible, and fully verifiable through automated grading. Here’s how we can help: * Design challenging repository-based systems engineering tasks * Build reference solutions, baselines, and adversarial submissions * Create deterministic automated graders with comprehensive hidden tests * Validate tasks through GitHub, CI, and iterative review Could you clarify which technical areas you want prioritized first, such as distributed systems, concurrency, protocols, or compilers?
$5 USD in 40 days
1.7
1.7

Hello, I’m a Senior Full-Stack Engineer with 8+ years of experience building and debugging complex software systems, backend services, APIs, databases, and AI-powered applications. I’m comfortable with Git/GitHub workflows, automated testing, debugging, system design, and writing clear technical specifications. This role is especially interesting to me because I enjoy breaking complex engineering problems into precise, testable requirements. I can create a complete task repository with the specification, reference implementation, automated grader, baseline solution, and adversarial tests to ensure the task is genuinely challenging and cannot be easily gamed. One task I would like to create is a distributed backend service involving concurrency, failure recovery, and deterministic behavior. The grader could validate correctness under concurrent workloads, recovery scenarios, and strict performance constraints while testing edge cases and potential reward-hacking approaches. I’m comfortable using AI tools to accelerate development and debugging while keeping the engineering decisions, validation, and task design fully my own. I’d be happy to contribute high-quality, rigorous engineering tasks to your evaluation pipeline. Kind regards, Juan
$5 USD in 40 days
0.0
0.0

Hello, I’m Adam, a Senior Software Engineer with 7+ years of experience building, debugging, and scaling complex software systems across backend, full-stack, and AI-powered applications. My experience includes building production systems involving: - Backend architecture and API development with Node.js, Python, FastAPI, Laravel, and PostgreSQL. - Distributed application design, data processing workflows, and system integrations. - AI-powered applications, LLM workflows, RAG systems, and evaluation-oriented development. - Writing clean technical specifications, debugging complex issues, and creating reliable test scenarios. - Git/GitHub workflows, CI/CD pipelines, Docker, and production deployments. I have strong experience using AI tools as part of the engineering workflow for code analysis, debugging, validation, and improving development efficiency, while maintaining independent technical judgment and system-level reasoning. I would be interested in contributing tasks based on areas such as backend reliability, API correctness, data consistency, AI application infrastructure, or production debugging scenarios. I’d appreciate the opportunity to discuss your evaluation framework and how my engineering background can contribute to creating challenging, high-quality tasks for AI coding agents. Thank you for your time and consideration. Best regards, Adam
$10 USD in 40 days
0.0
0.0

Hello, I am a senior software engineer with over ten years of experience designing and building scalable backend systems, distributed services, and full‑stack applications. I have deep expertise in Java, Node.js, Python, and cloud platforms, and I regularly write clear technical specifications, API contracts, and test suites as part of my development workflow. My work with microservices, concurrency‑safe services, and CI/CD pipelines has given me strong practice in creating reproducible repositories, reference implementations, and automated validation checks—exactly the skills needed to craft Frontier‑style tasks, their public instructions, reference solutions, and rigorous graders. I am comfortable working independently, iterating based on feedback, and using Git/GitHub for collaborative review. I would welcome the opportunity to turn real‑world systems challenges into verifiable coding tasks for AI agent evaluation. Best regards
$50 USD in 40 days
0.0
0.0

atlanta, United States
Payment method verified
Member since Oct 24, 2019
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$10-30 USD
$10-30 USD
$15-25 USD / hour
$10-50 USD
$250-750 AUD
$50-100 USD
₹250000-500000 INR
$15-25 USD / hour
$250-750 USD
min $50 USD / hour
min $50 USD / hour
$5000-10000 USD
$30-250 USD
$250-750 USD
€8-30 EUR
$1500-3000 USD
$250-750 AUD
$30-250 CAD
₹1500-12500 INR
$5000-10000 USD
$10-20 SGD / hour
$8-15 USD / hour