All vacancies
Python Engineer, AI Coding Agent Evaluator
G2i
Remote · worldwide, UTC+4 acceptedSalary not disclosedcontractVerified 4 days agoHimalayas
We’re looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not a traditional engineering role.
Responsibilities
- Evaluate AI-generated coding interactions end-to-end
- Judge whether outputs are
- Correct (at a high level)
- Aligned with how a strong engineer would think
- Assess the quality of explanations and reasoning, not just code
- Distinguish between different levels of response quality (e.g. what makes something a 2 vs 4)
- Provide clear, opinionated feedback on
- What felt “off” or misleading
- Help define what great looks like when interacting with tools like Cursor
- What We Mean by “Taste”
Languages
- Work format
- Remote
- Seniority
- Mid
- Posted
- 10 Sept 2026 (1w ago)
- Last verified
- 13 Sept 2026
- Apply by
- 9 Nov 2026
