Applied Computer Science Benchmark Specialist
Role Overview
We are seeking expert computer scientists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core computer science domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.
You will be assigned one of two task types
- Question Authoring, Create original, challenging multiple-choice questions in your area of computer science expertise, rate their difficulty, and submit them for review.
- Question Verification, Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made.
Computer Science Domains Covered
Accelerator / GPU Kernel Engineering, Formal Methods & Automated Reasoning, Computer Architecture & Accelerators, Distributed Systems, DevOps & Site Reliability, Data Engineering & Databases, Cloud & Infrastructure, OS & Systems Kernel, Machine Learning Engineering, Web & API Development, Embedded Systems Engineering, Computer Graphics & Game Development, Mobile Engineering.
Key Responsibilities
- Author original computer science questions that test deep conceptual understanding, not surface-level recall
- Ensure questions are unambiguous, self-contained, and precisely defined, all necessary information must be in the problem statement
- Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above)
- Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers
- Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format
- Supply 1, 5 academic references per question from reputable sources (peer-reviewed journals, university repositories)
- For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made
Ideal Qualifications
- PhD or doctoral candidate in Computer Science, Electrical Engineering, or a closely related field
- Master's degree considered for candidates with exceptional depth in a specific subdomain
- Strong command of graduate-level CS theory, algorithms, systems design, and/or machine learning
- Research publications, industry experience at top tech companies, or competitive programming background is a strong plus
- Excellent written English and ability to express complex ideas clearly and concisely
More About the Opportunity
- Expected commitment: 10+ hours/week
- Asynchronous, fully remote work
Details
- Pay: $66 to $84 per hour (hourly)
- Commitment: Flexible • remote
- Eligible locations: United States
- Platform: Mercor (weekly payouts via Stripe or Wise)
Submit your application via the link below. Qualified candidates move through Mercor's short AI interview and selection process. Projects can be extended, shortened, or concluded early depending on needs and performance.
Stay Updated on Roles Like This
Subscribe to receive fresh openings aligned with AI Training & Evaluation and AI training roles across Mercor and JobHub by NeonLabs