logo
Gold Group Ltd

Software Engineer

North America

July 2, 2026 at 4:00 PM

✨ AI summary

Location

  • United States (Remote)

Languages

  • Python

Frameworks

  • Inspect (bonus)

Software Engineer – AI Benchmarking & LLM Evaluations

Remote


I'm currently working with a leading AI research institute on a fantastic opportunity for a Software Engineer to join their Benchmarking team.


If you're fascinated by frontier AI models and enjoy understanding how they perform rather than simply building applications with them, this could be an excellent fit.


You'll help develop and run evaluations of the latest AI models, build new benchmarks, maintain evaluation infrastructure, and collaborate directly with researchers producing work that influences policymakers, industry leaders, and the wider AI community.


What they're looking for:

  • Strong software engineering experience (language agnostic – Python preferred)
  • An interest in LLM evaluations, benchmarking, or AI capability testing
  • Curiosity about frontier AI and a research-oriented mindset
  • Someone who enjoys experimentation, solving difficult technical problems, and improving evaluation frameworks


Experience with evaluation frameworks such as Inspect is a bonus but certainly not essential.

This is an opportunity to work on research that has a genuine impact on how the world understands AI progress, rather than building commercial products.


  • Fully remote
  • Three international company retreats each year
  • Flexible working hours
  • Open to candidates across many countries (with a preference for US or European time zones)