Data Science Evaluator (contract)
Overview
MarketerHire is seeking a Data Science Evaluator for a 12-week independent contractor engagement focused on evaluating AI systems performing real-world data science work. The role requires 2–3 years of professional data science experience, strong Python, SQL, statistics, machine learning, experimentation, and causal inference skills, plus excellent written English and the ability to apply consistent judgment through ambiguity.
Responsibilities
<h2>Key Responsibilities</h2><ul><li>Design task-specific grading criteria for exploratory analyses, statistical models, ML pipelines, A/B test write-ups, and technical notebooks.</li><li>Evaluate AI-generated and human-created data science work against established grading criteria.</li><li>Write clear, evidence-based justifications for every evaluation score.</li><li>Apply consistent evaluation judgment across varied technical deliverables and use cases.</li><li>Reason through ambiguity when assessing the quality, correctness, and completeness of submitted work.</li><li>Identify subtle errors in polished data science analyses and technical documentation.</li><li>Clearly separate established knowledge from inferences when documenting evaluation conclusions.</li><li>Iterate on evaluation work based on feedback from senior reviewers.</li><li>Follow complex instructions accurately while maintaining consistency across assigned evaluations.</li><li>Learn new tools quickly and work independently throughout the contract engagement.</li></ul>
Requirements
<h2>Required Qualifications</h2><ul><li>2–3 years of professional data science experience.</li><li>Strong skills in Python, SQL, statistics, machine learning, experimentation, and causal inference.</li><li>Experience with exploratory analysis, statistical modeling, ML pipelines, A/B testing, and technical notebooks.</li><li>Excellent written English.</li><li>Ability to follow complex instructions, reason through ambiguity, and identify subtle errors in polished work.</li><li>Ability to maintain a clear separation between established knowledge and inferences.</li><li>Ability to work independently, learn new tools quickly, and take feedback well.</li><li>Ability to create grading criteria, evaluate AI output, and review technical work.</li><li>Based in and working from the US, Canada, or UK.</li><li>Available for 30–40+ hours per week for the full 12-week engagement.</li></ul>
Benefits
<h2>Benefits and Perks</h2><ul><li>Fully remote contract engagement.</li><li>Competitive hourly rate of $100–$150.</li><li>Flexible scheduling within the required 30–40+ weekly hours.</li><li>Weekly payment via Stripe or Wise.</li><li>Potential for projects to be extended based on business needs and performance.</li><li>Opportunity to contribute to the evaluation of real-world AI data science capabilities for leading AI research organizations.</li></ul>