JobTarget Logo

Senior Software Engineer – LLM Evaluation in Haciendas del Canada, Nuevo León at Jobgether

NewJob Function: Information Technology
Jobgether
Haciendas del Canada, Nuevo León, 66054, Mexico
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Senior Software Engineer LLM Evaluation

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer – LLM Evaluation based in Canada.

This part-time consulting opportunity is designed for experienced software engineers interested in shaping the evaluation of advanced large language models.
You will curate high-quality code, develop software solutions, and assess AI-generated implementations against professional engineering standards.
The role spans multiple programming languages and covers the full software-development lifecycle, from architecture and prototyping to deployment, monitoring, and maintenance.
You will also help create automated verification mechanisms and benchmarks that make model evaluation more rigorous and reproducible.
Working alongside research and technical teams, you will help identify model strengths, weaknesses, and recurring coding errors.
Your engineering judgment will directly contribute to improving coding-focused AI evaluation systems and research.
The engagement is fully remote and flexible, with a minimum commitment of 10 hours per week and the potential to work up to 40 hours weekly.

Accountabilities:
  • Curate high-quality software examples and datasets for model training, coding benchmarks, and technical evaluation.

  • Develop precise solutions to software-engineering tasks and correct or improve implementations across multiple programming languages.

  • Evaluate AI-generated code for technical correctness, maintainability, efficiency, scalability, reliability, and alignment with professional engineering standards.

  • Identify implementation weaknesses, recurring model-generated errors, and patterns that reveal gaps in AI coding capabilities.

  • Provide clear, structured rationales supporting technical evaluation decisions.

  • Build agents and automated mechanisms capable of assessing code quality and verifying software solutions consistently.

  • Develop reliable checks that enable reproducible evaluation across repeated engineering assignments.

  • Assess model capabilities across the complete software-development lifecycle, including prototyping, architecture, API design, implementation, experimentation, launch, monitoring, and maintenance.

  • Review AI reasoning and technical decisions against realistic production engineering expectations.

  • Collaborate with research and cross-functional technical teams to define evaluation strategies and improve coding-focused benchmarks.

  • Contribute to datasets used for training and benchmarking while maintaining rigorous standards for technical accuracy and quality.

  • Support iterative improvements to evaluation systems by comparing model performance against real-world software-engineering practices.

  • Complete all project work without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any third party.

Requirements:
  • 3+ years of professional software-engineering experience.

  • Strong full-stack development capabilities and experience building scalable, production-grade software.

  • Strong understanding of software architecture, system design, API design, and production implementation.

  • Deep knowledge of software development, debugging, code review, and code-quality assessment.

  • Demonstrated experience reviewing, troubleshooting, and improving complex software implementations.

  • Proficiency in one or more relevant programming languages, including Python, JavaScript, Java, C++, Rust, or similar technologies.

  • Experience with ReactJS, C, Go, or additional programming languages is valuable depending on project requirements.

  • Familiarity with software monitoring, operational maintenance, and production reliability.

  • Ability to reason across the full software-engineering lifecycle and assess technical decisions from development through ongoing operation.

  • Strong analytical and problem-solving abilities, with a rigorous and detail-oriented approach to technical evaluation.

  • Excellent written and verbal communication skills, including the ability to produce concise, well-structured evaluation rationales.

  • Ability to distinguish between technically correct solutions and implementations that may introduce scalability, reliability, maintainability, or architectural concerns.

  • Comfortable working independently and collaborating remotely with research, engineering, and other technical teams.

  • Based in Canada, the United States, or an eligible Western European country.

  • Willingness to complete a required AI video interview as part of the application process.

Benefits:
  • Fully remote, part-time independent contractor engagement.

  • Flexible workload starting at 10 hours per week, with the potential to work up to 40 hours per week.

  • Approximately one-month initial project duration, with potential extension based on performance and project fit.

  • Opportunity to contribute directly to advanced LLM evaluation, coding benchmarks, and AI-assisted software-engineering research.

  • Exposure to cutting-edge AI evaluation workflows and research involving real-world software-development scenarios.

  • Opportunity to apply professional engineering expertise to improve how AI systems are evaluated across the software lifecycle.

  • Flexible consulting structure suited to experienced engineers seeking project-based work.

  • International remote collaboration with research and technical professionals.

  • No medical insurance or paid-leave benefits are included under the independent contractor arrangement.

  • Compensation is project-specific and was not specified in the available job materials.

How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

Haciendas del Canada, Nuevo León, 66054, Mexico

Frequently asked questions about this position

Similar Jobs In Haciendas del Canada, Nuevo León

New

Forward Deployed Engineer (AI Capability Center)

Jobgether
Haciendas del Canada, Nuevo León

Senior Infrastructure Engineer

Jobgether
Haciendas del Canada, Nuevo León
New

Dynamics 365 Customer Engagement (CE) Technical Architect

Jobgether
Haciendas del Canada, Nuevo León
New

Software Engineer (Front and Back end)

Jobgether
Haciendas del Canada, Nuevo León
New

Senior Software Engineer, Site Reliability Engineering

Jobgether
Haciendas del Canada, Nuevo León