Senior Python Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr) in New York at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Python Engineer - AI Code Evaluation based in the United States.
This is a senior-level contract opportunity focused on evaluating the quality of interactions between developers and modern AI coding agents.
Rather than building production software, you’ll apply your engineering judgment to assess whether AI-generated responses reflect strong technical reasoning and practical engineering standards.
You’ll review coding-agent interactions for usefulness, correctness, clarity, and overall quality of engineering judgment.
The role requires a strong sense of engineering “taste” and the ability to make subjective assessments that are still rigorous and well supported.
You’ll help distinguish excellent AI-assisted development experiences from responses that are technically valid but confusing, misleading, or poorly reasoned.
You’ll work with tools such as Codex, Claude Code, and Cursor in a flexible, remote environment.
The engagement offers senior engineers the opportunity to directly influence how AI coding systems are evaluated and improved.
- Evaluate AI coding interactions: Review AI-generated coding sessions end to end and assess whether responses are useful, correct at a high level, and aligned with how an experienced engineer would approach the problem.
- Assess engineering judgment: Evaluate whether AI agents demonstrate sound reasoning, appropriate trade-offs, clear decision-making, and practical software-engineering judgment rather than simply producing syntactically correct code.
- Review explanations and reasoning: Judge the quality of model preambles, explanations, reasoning, and guidance, identifying whether they genuinely help developers or simply generate technically accurate but unhelpful output.
- Provide structured feedback: Clearly and directly explain what worked, what did not, and what felt misleading, ineffective, or inconsistent with strong engineering practice.
- Differentiate response quality: Apply consistent standards to distinguish varying levels of AI performance and articulate what separates an average response from an excellent one.
- Help define evaluation standards: Contribute to the evolving understanding of what high-quality AI-assisted software development should look like, including interactions with AI-first development environments.
- Apply independent judgment: Make subjective but rigorous assessments based on extensive engineering experience, without needing to execute or deeply inspect every line of generated code.
- Senior engineering experience: Staff-, Principal-, or similarly experienced software engineering background, with a strong track record of making technical decisions and maintaining a high standard for software quality.
- Strong programming expertise: Deep hands-on experience with Python and/or TypeScript/JavaScript, with the ability to quickly understand and assess modern software implementations.
- AI coding experience: Practical experience using AI-assisted development tools such as OpenAI Codex, Claude Code, Cursor, or comparable coding agents.
- Modern development knowledge: Strong familiarity with contemporary AI-assisted software-development workflows and how developers interact with coding agents.
- Engineering judgment: Ability to recognize whether an AI response reflects strong engineering thinking, including useful explanations, sound trade-offs, appropriate guidance, and trustworthy recommendations.
- Critical evaluation skills: Comfortable evaluating code and technical reasoning at a high level without needing to execute or inspect every implementation detail.
- Communication skills: Able to provide concise, direct, and opinionated feedback that clearly explains why an interaction succeeds or falls short.
- High quality bar: Strong personal standards for what constitutes effective, maintainable, and professional engineering.
- Preferred experience: Exposure to AI-first IDEs, prompt design, model evaluation, evaluation workflows, or engineering-standard development is a plus.
- Additional advantage: Experience mentoring senior engineers or helping establish engineering standards and best practices is beneficial.
- Compensation: $100–$200 per hour, depending on experience and fit.
- Flexible engagement: Approximately 10–20 hours per week.
- Remote work: Fully remote contract engagement with flexibility across eligible locations.
- Short-term focused engagement: Initial engagement runs through early May, with potential for extension.
- Fast start: Opportunity to begin as soon as possible.
- Meaningful AI impact: Directly contribute to evaluating and improving how AI coding agents support professional software engineers.
- Senior-level scope: Apply your engineering expertise to challenging questions around AI reasoning, developer experience, and software-engineering quality.
- Streamlined selection: Hiring process consists of a take-home evaluation exercise followed by one behavioral interview.