Loading opportunity…
AI Training · hourly
Evaluate the quality, correctness, and reproducibility of software engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric based written feedback.…
THE OPPORTUNITY
Evaluate the quality, correctness, and reproducibility of software engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric based written feedback.…