
Closed
Posted
Software Engineering Task Author (Contract) About the Role We're looking for an experienced software engineer to design and build Frontier-style software engineering tasks — long-horizon, repository-based coding challenges used to train and evaluate AI coding agents. You'll take real engineering problems from your own area of expertise and turn them into rigorous, verifiable tasks: a seeded repository, a clear public specification, and an automated grader that checks observable behavior (builds, hidden test suites, protocol conformance, concurrency/recovery scenarios, deterministic replay, numeric budgets, etc.). This is deep, research-grade work — closer to writing a hard systems/algorithms exam with an autograder than to typical QA or content work. What You'll Do Design non-trivial engineering problems grounded in real-world systems (e.g., concurrency bugs, recovery/consistency issues, legacy modernization, protocol implementation, performance-constrained rewrites). Build a starter repository, write a complete public specification ([login to view URL]), and implement a reference solution that solves the task correctly. Write automated grading code that measures exactly what the instructions ask for — no hidden requirements, no LLM judges, no shortcuts. Produce a no-op/naive baseline to confirm the task has real difficulty, and adversarial "attack" submissions to confirm the grader can't be gamed. Validate, iterate, and submit tasks via pull request through a structured review pipeline. Use AI tools to accelerate parts of your workflow (explaining code, debugging validation errors, reviewing your own logic) — but the engineering judgment and task design must be your own. What We're Looking For Strong professional software engineering background (backend systems, distributed systems, compilers, low-level/systems programming, or similar) — ideally in an area where you have real depth and war stories, not just familiarity. Comfort with concurrency, correctness, and testing — the ability to reason precisely about what a grader should and shouldn't check. Experience writing clear technical specifications and/or authoring hard technical interview problems, coding challenges, or test suites. Git/GitHub workflow fluency (branches, PRs, CI). Self-directed: comfortable working from written guidelines, validating your own work, and iterating based on automated feedback and review. Bonus: experience with reinforcement learning, model evaluation, or "reward hacking" adversarial thinking (i.e., trying to break your own grader before someone else does). Nice to Have A specific technical niche you know cold (e.g., database internals, network protocols, embedded systems, compilers/language runtimes, numerical computing) that would make for a compelling, hard-to-fake task. Prior experience creating technical assessments, benchmarks, or eval datasets. Engagement Details Contract / freelance, task-based. Deliverables are reviewed against defined structure, CI, and quality checks before acceptance. Ongoing support and clarification provided via a shared Discord community. To apply, please share examples of complex systems you've built or debugged, any prior experience authoring technical assessments or benchmarks, and a brief note on what kind of engineering problem you'd want to turn into a task.
Project ID: 40672674
35 proposals
Remote project
Active 3 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
35 freelancers are bidding on average $14 USD/hour for this job

★•══•★ Hi client ★•══•★ Got it. I've read through what you need, and it's right up my alley. I don't just patch things up—I dig in, find the root cause, and make sure it's solid before handing it back. No surprises later, just clean, working results. You'll get clear communication and a straightforward approach. No tech jargon overload, just honest work. Shall we get started? I'm ready when you are. Best regards, Rico
$8 USD in 40 days
5.0
5.0

Hi there, Your task pipeline needs someone who can turn real backend problems into rigorous, testable engineering challenges. I have strong expertise in Backend Development, Java, and Software Development, and I’d approach this by designing a seeded repository, writing a precise public spec, and building an automated grader with Software Testing that checks only observable behavior. I’m also comfortable using Git workflows to keep the task, reference solution, and validation steps clean and reviewable. Best regards, Ian
$20 USD in 21 days
4.2
4.2

Having built AI systems for real-world deployment, I know firsthand the value of rigorous and relevant software engineering tasks. My background in AI development, backend work and Git -- combined with my experience using a range of technologies from React to MQTT -- aligns closely with what you're looking for. On the one hand, my familiarity with systems such as Odoo ERP and embedded hardware found in MQTT-connected sensor networks perfectly positions me to design intricate tasks tying together distinct technical niches. Often, prototypes fall short when deployed in actual workflows; this is why I prioritize production infrastructure. My ability to "walk-the-talk" in areas such as distributed systems, performance-constrained rewrites, and low-level/systems programming means I can offer you a rare mix of in-depth expertise sprinkled with war stories that cannot be learned from mere familiarity or documentation. Not only do I assure you my work will have real impact and difficulty but also demonstrate the tenacity of being self-directed worker who needs minimal guidance and can iterate effectively based on automated feedback and review. With me onboard, rest assured your tasks will be no closer to typical QA or content work but rather a deep dive into designing challenging evolutions sessions for AI coding agents using research-grade methods tailored by an experienced software engineer like me!
$5 USD in 40 days
4.3
4.3

⮞⮞⮞⮞⮞ Dear client ⮜⮜⮜⮜⮜, Thank you for seeing my proposal. I have carefully reviewed your Software Engineering Task Author role and I am confident I can design rigorous, research-grade engineering tasks. -What are you looking to solve in this project? You need an experienced software engineer to design and build Frontier-style engineering tasks—long-horizon, repository-based coding challenges with seeded repositories, clear specifications, and automated graders that verify correctness through builds, test suites, and performance constraints. -What I can do for you in this project? I have deep experience in backend systems, concurrency, and correctness. I can design non-trivial tasks, write complete specs, implement reference solutions and automated graders, and create adversarial test cases to ensure robustness. I enjoy breaking my own graders before someone else does. ⚠️ If you want to solve more, I will do—advanced grading logic, multi-phase tasks, or performance-constrained challenges. Thank you for your time. I am confident I can deliver rigorous, hard-to-game engineering tasks. Best Regards
$5 USD in 60 days
1.8
1.8

✅✅Hi, there✅✅ I am excited about the opportunity to design and build rigorous software engineering tasks, focusing on real-world systems and ensuring precise grading without hidden requirements. This project aligns perfectly with my strong background in backend systems and experience in authoring technical challenges. I understand that you're looking for someone to create non-trivial engineering problems, develop public specifications, and implement automated grading systems. I can deliver high-quality tasks that reflect genuine complexity and validate real engineering scenarios. Moreover, I can ensure that the tasks are well-documented and structured, providing clear instructions and feedback mechanisms. I’m eager to connect and discuss how my skills can contribute to your project, as I believe collaboration will enhance the outcomes we aim for. Looking forward to hearing from you! Best, Predrag
$5 USD in 40 days
1.9
1.9

I am excited about the opportunity to contribute as a Software Engineering Task Author for Frontier-style challenges. With a strong background in backend systems and distributed systems, I excel at designing complex engineering problems that precisely test concurrency, correctness, and system behavior. I am proficient in Git/GitHub workflows and experienced in writing clear, technical specifications and coding challenges. My approach emphasizes rigorous, research-grade task design with robust automated grading and adversarial testing to prevent shortcut solutions. I look forward to leveraging my expertise to create meaningful, real-world problem sets for AI coding agents. Thank you for considering my application. Best regards, David.
$20 USD in 34 days
0.0
0.0

The hardest part of what you're describing isn't writing the code — it's designing a grader that can't be gamed, and that's exactly where most task authors fall short. My approach: I'd anchor each task in a real system failure I've personally debugged — concurrency races, consistency edge cases, protocol misimplementations — then build the spec, starter repo, reference solution, and grader as a tight unit where the observable behavior is the only thing that matters. No LLM judges, no vibes-based scoring. A few things I'd bring to this specifically: I think adversarially from the start, meaning I'm already trying to break my own grader while I'm writing it, and I have experience writing technical specs that are precise enough to be unambiguous but not so prescriptive that they hand the solution away. Two quick questions before I dive in: Is there a preferred domain you're currently under-indexed on (e.g., you have plenty of concurrency tasks but need more compiler or network protocol work)? And does the $30 rate apply per accepted task after review, or is there a partial rate for tasks that need significant revision cycles? Happy to share a sample task outline or a system I'd draw from — just say the word.
$5 USD in 40 days
0.0
0.0

Hi Designing rigorous, Frontier style tasks with repository-based coding challenges aligns with my expertise in AI-powered applications and software engineering. The core technical challenge here is developing tasks that simulate real world engineering problems with automated grading. This involves crafting a seeded repository, a detailed public specification, and implementing a robust automated grading system, all built to stress actual engineering skills. Having developed intelligent applications and scalable web platforms using technologies like Python, JavaScript, and cloud services, I understand the complexity of creating tasks that are not only challenging but also verifiable. My experience with AI driven workflow automation includes building systems that evaluate complex scenarios, akin to what's needed here for grading. I appreciate the emphasis on using AI tools for workflow acceleration but keeping the engineering judgment authentic, something I always value in AI development. I'm ready to design these tasks and walk you through a detailed approach to ensure these challenges both test and teach effectively. Let's dive deeper into how I can assist you. Thanks, Oswaldo
$50 USD in 40 days
0.0
0.0

Hi, I’ll focus on repository-based tasks with deterministic graders, adversarial tests, and clearly measurable acceptance criteria. Creating challenging engineering tasks that expose real coding-agent limitations without hidden requirements or gameable grading is the goal. I’ve worked on projects where backend correctness, concurrency, recovery behavior, Git workflows, automated testing, and precise technical specifications required careful reasoning beyond ordinary feature development. I’d like to win this project and I’m confident I can deliver high-quality tasks through your review and CI pipeline if awarded. I’d particularly explore concurrency and distributed-state problems where naive implementations appear correct but fail under recovery, ordering, or deterministic replay, then validate each grader against baseline and adversarial submissions. Thanks.
$5 USD in 40 days
0.0
0.0

✌️Hi, there I can create rigorous, verifiable coding tasks using my deep experience in designing complex backend systems and automating grading processes. This project aligns perfectly with my skills, and I'm excited to bring real-world engineering problems to life with 100% accuracy. You need someone who can design challenging tasks like concurrency bugs and protocol implementation while providing clear specifications and automated grading. I understand that the work requires a strong grasp of software engineering principles, as well as the ability to validate and iterate through a structured review. By leveraging my background, I can ensure that the tasks I create are not only challenging but also educational, fostering a better understanding of engineering concepts. I look forward to connecting with you soon, as collaboration will be key to meeting your deadlines. Best, Filip
$5 USD in 40 days
0.0
0.0

Hi, ⚡ I guarantee a 100% success rate.⚡ I understand this role needs someone who can turn real engineering problems into difficult, fair and automatically graded coding tasks for AI agents. My experience with full stack development, AI systems, backend architecture, testing and debugging gives me a strong foundation for this work. I would create a realistic repository, precise instructions, reference solution, automated tests and adversarial submissions, then validate that the grader checks behavior rather than superficial implementation. I would be interested in creating a task around concurrency, API reliability and recovery where agents must handle failures without breaking data consistency. Thank you.
$5 USD in 40 days
0.0
0.0

Hello, I am a seasoned AI Engineer specializing in creating AI-powered workflows and automations. With over 12 years of experience in technology solutions, I am well-equipped to design and build Frontier-style software engineering tasks for training and evaluating AI coding agents. My expertise lies in crafting non-trivial engineering problems grounded in real-world systems, designing clear technical specifications, and implementing rigorous automated grading systems. I have a strong background in backend systems, distributed systems, and compilers, enabling me to excel in this role. I am self-directed, detail-oriented, and proficient in Git/GitHub workflows. I am excited about the opportunity to contribute to this project and look forward to discussing further details. Best regards, Noman
$2 USD in 40 days
0.0
0.0

With over a decade of professional software engineering under my belt, I can confidently say that I am adept at handling the challenges outlined in this position. My diverse range of experience from backend to distributed systems will be invaluable in designing and building substantial coding tasks that draw on real-life scenarios. Furthermore, I have managed numerous projects with concurrent processes, securely ensuring system correctness while maintaining high-performance levels. In terms of technical specifications, as a Senior Full-Stack Architect, clear and concise communication is paramount. Whether it's coding challenges or test suites, I have regularly crafted rigorous technical instructions that importantly do not hide any requirements. Additionally, my proficiency in Git/GitHub aligns with your workflow fluency requirement, which ensures smooth navigation through branching, Pull Requests, and Continuous Integration. What sets me apart from the competition is my appreciation for the intelligent architecture behind any system. This allows me to develop clever baseline submissions to assess task difficulty and parallelly erect adversarial "attack" solutions - a process where I get a real kick out of trying to game my own grading system (reward-hacking). This knack for critically evaluating a system will ensure that the tasks I create are both challenging for AI coding agents and incorruptible. Please have a look on my profile. Regards, Firasat W
$5 USD in 40 days
0.0
0.0

With a 10+ year career in full-stack development, my proficiency in Java and experience designing, developing, and scaling web, mobile, cloud, and AI-powered applications is a strong fit for your project. In particular, my expertise in backend systems, distributed systems, low-level/systems programming and related areas would be especially value-adding given the nature of the tasks you need created. In addition to my software engineering skills, I also have extensive knowledge and comfort with concurrent code implementation and testing. I am both familiar and experienced with working in Git/GitHub workflows and understand how crucial it is when developing complex projects collaboratively. More importantly, I can bring to bear a unique mix of technical knowledge and "reinforcement learning" thinking that aligns well with your project needs - which involves creating challenging coding tasks that cannot be easily gamed. My familiarity with creating technical assessments and benchmarks could hugely benefit your project by ensuring the rigor, verifiability, and complexity of the tasks I create meet your expectations. I am attracted to this opportunity not just because it aligns well with my technical capabilities but also because it provides an avenue to apply my passion for building high-quality software efficiently to improve overall product experience. I look forward to bringing this excitement for meaningful projects to the team!
$5 USD in 40 days
0.0
0.0

Hi, I’ve worked on backend and engineering problems where correctness matters more than simply making the code pass. Your Frontier-style tasks need strong specs, reproducible repos, reference solutions, and graders that catch real behavior rather than easy hacks. I’d build each task around a genuine systems problem, with clear acceptance criteria, deterministic tests, edge cases, and adversarial submissions to verify the grader cannot be gamed. I’m comfortable with Git/PR workflows, debugging, test automation, and using AI tools without outsourcing the engineering judgment. One task idea I’d enjoy: a concurrent job system with retries, crash recovery, ordering guarantees, and strict resource limits, where hidden tests verify correctness under race conditions. What languages and repository patterns does your current benchmark favor? I can also propose the first task structure and grading strategy.
$5 USD in 40 days
0.0
0.0

I noticed that you are prioritizing "adversarial" submissions to test grader robustness—have you already identified the specific edge cases where an LLM agent is most likely to "reward hack" your current reference solutions, or are you looking for an engineer to stress-test those boundaries from scratch? I would focus on creating isolated, reproducible environments that account for concurrency and state-consistency issues, ensuring the autograder is truly immutable against "shortcuts." If you have a draft repository or a specific problem domain you are starting with, I would be happy to review the current specification and suggest where we might tighten the grading logic.
$10 USD in 40 days
0.0
0.0

Ethan here, from South Africa. Your project immediately caught my eye. I'm really excited to partner with you. I see you need someone to design rigorous, verifiable Frontier-style software engineering tasks that address real-world systems. I would approach this by focusing on creating clear specifications and starter repositories that accurately reflect the complexity of the engineering challenges. By utilizing AI tools for parts of the workflow, I can enhance the process while ensuring that the core design and engineering judgment remain my own. I would ensure quality by validating and iterating on tasks through a structured review pipeline, making certain that each task meets the criteria without hidden requirements. What you truly need is not just tasks, but well-designed engineering challenges that assess real skills and knowledge. Please feel free to reach out so we can connect and further explore how I can contribute to your project's success. Kind regards, Ethan
$3 USD in 8 days
0.0
0.0

Hi, I’ve reviewed your project and understand that you need help with designing and building Frontier-style software engineering tasks. I can assist you with creating rigorous, verifiable coding challenges grounded in real-world systems, ensuring that each task has a clear specification and an automated grading system. My experience in backend systems and distributed systems will allow me to handle the complexities of concurrency, correctness, and testing efficiently while keeping communication clear throughout the project. I can also help with writing clear technical specifications and iterating on tasks based on feedback, ensuring the final result meets your expectations. I'd love to chat about your project! The worst that can happen is you walk away with a free consultation. Regards, JaniceR92
$4 USD in 7 days
0.0
0.0

Just check me out and confirm a try will convince you I don't need to put too much work because I know my work thanks
$200 USD in 10 days
0.0
0.0

Hi, Concurrency and state-recovery edge cases are the real risk when a Frontier-style task is graded automatically; subtle nondeterminism and flakey test harnesses create false positives that waste reviewer time. I have 12+ years building backend systems and debugging production concurrency, recovery, and correctness issues in .NET Core microservices and messaging pipelines. I wrote an automated fault-injection test harness for a distributed job scheduler that reproduced a rare race condition and guided a fix rolled into production. I have authored technical exercises used in senior hiring loops that included seeded repos, clear public specifications, and CI graders (unit and property tests) to separate correct implementations from brittle workarounds. I will start by drafting a minimal seeded repository and a public specification that makes failure modes explicit, then implement a reference solution and a grader that uses deterministic fixtures and fuzzed adversarial submissions to validate difficulty. - Examples of complex systems you've built or debugged I built and operated a .NET Core microservice mesh with eventual-consistency recovery, and I debugged a concurrency bug in a distributed scheduler using targeted trace logging and deterministic replay to isolate the race. - Any prior experience authoring technical assessments or benchmarks I authored senior-level coding assessments with seeded repos, automated graders, and naive/adversarial submissions for interview and benchmarking purposes in prior hiring cycles. - A brief note on what kind of engineering problem you'd want to turn into a Frontier-style task I would turn a real-world recovery/correctness problem into a task: implement a durable message deduplication layer for a concurrent consumer group, with injected network partitions and time-based clock drift to test correctness under adversarial conditions. - Would you prefer tasks focused on concurrency correctness, or on fault-recovery under partitions? - Do you have a preferred language/runtime and CI environment for submissions (for example Node, Python, or .NET Core)? Would love to jump on a quick call if that works for you. Christopher
$2 USD in 1 day
0.0
0.0

atlanta, United States
Payment method verified
Member since Oct 24, 2019
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$30-250 USD
$750-1500 USD
$10-30 USD
₹12500-37500 INR
₹1500-12500 INR
$15-25 USD / hour
$10000-20000 USD
$8-10 USD / hour
₹1500-12500 INR
$250-750 AUD
$25-50 USD / hour
$30-250 CAD
$25-50 AUD / hour
₹600-1500 INR
min $50 USD / hour
$250-750 USD
₹600-1500 INR
$15-25 USD / hour
$30-250 USD
$10-20 SGD / hour