Research Engineer - AI Evaluation & Benchmark Development
Job Description
Hiring: Research Engineer – AI Evaluation & Benchmark Development (USA | Remote)
\n
We are seeking exceptional Research Engineers, Applied Scientists, AI Researchers, and Machine Learning Researchers to join a cutting-edge AI evaluation initiative focused on developing next-generation research engineering benchmarks for frontier AI systems.
\n
\n
This opportunity is ideal for candidates with a strong research background who enjoy solving complex technical problems, designing experiments, and contributing to AI evaluation and benchmarking.
\n
\n
Target Domains
\n
We are specifically looking for candidates with expertise in one or more of the following domains:
\n
Computer Science & Software Engineering
\n
- \n
- Python Programming
- Software Development
- Git & Development Infrastructure
- AI Coding Assistants / Agentic Workflows
- Software Quality & Debugging
\n
\n
\n
\n
\n
\n
\n
Machine Learning & Artificial Intelligence
\n
- \n
- Machine Learning
- Deep Learning
- Large Language Models (LLMs)
- Reinforcement Learning
- ML Experimentation
- Model Training & Evaluation
- AI Evaluation
\n
\n
\n
\n
\n
\n
\n
\n
\n
Data Science & Quantitative Analysis
\n
- \n
- Data Science
- Statistical Analysis
- Data Analytics
- Experimental Data Analysis
- Jupyter Notebook / Google Colab
\n
\n
\n
\n
\n
\n
\n
STEM Research & Experimental Methodology
\n
- \n
- Research Engineering
- Scientific Computing
- Experimental Design
- Hypothesis Testing
- Computational Sciences
- Computational Mathematics
- Computational Physics
- Computational Biology
- Computational Social Sciences
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Preferred Experience
\n
- \n
- AI Safety
- LLM Red Teaming
- Benchmark Development
- Test Engineering
- Quality Assurance
- AI System Evaluation
\n
\n
\n
\n
\n
\n
\n
\n
Key Responsibilities
\n
- \n
- Design and develop complex research engineering tasks that simulate real-world AI and machine learning challenges.
- Create problem statements and benchmark scenarios for evaluating advanced AI systems.
- Design, execute, and analyze machine learning experiments.
- Evaluate AI model performance and document experimental findings.
- Collaborate with research teams to improve benchmark quality and evaluation methodologies.
\n
\n
\n
\n
\n
\n
\n
Minimum Qualifications
\n
- \n
- Master's or PhD in Computer Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics, Physics, Data Science, Computational Sciences, or another STEM discipline.
- Minimum 1–2 years of experience in research or research engineering.
- Strong Python programming skills.
- Hands-on experience with Machine Learning, LLMs, experimentation, and data analysis.
- Proficiency with Git, IDEs, Jupyter Notebook, or Google Colab.
- Strong analytical thinking, scientific methodology, and problem-solving abilities.
- Excellent written and verbal communication skills in English.
- Candidates from top-tier universities are highly preferred
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Preferred Qualifications
\n
- \n
- Experience working alongside Research Scientists or AI Research teams.
- Publications in AI, Machine Learning, Data Science, or related research areas.
- Experience with AI benchmarking, LLM evaluation, or model testing.
- Familiarity with agentic AI systems and research workflows.
\n
\n
\n
\n
\n
\n
Work Details
\n
Location: Remote (United States)
\n
Availability: 35 hours per week
\n
Schedule: Monday to Friday
\n
\n
Working Hours: 7 hours per day between 6:00 AM and 6:00 PM PST
\n
If you have a passion for research, experimentation, and advancing AI through rigorous evaluation and benchmarking, we would be pleased to hear from you.
\n
\n
To apply, please send your updated resume to sajid.ahmed@truelancer.com or connect with me on LinkedIn for additional information.
\n
\n
#Hiring #ResearchEngineer #ResearchScientist #AppliedScientist #ArtificialIntelligence #MachineLearning #LLM #DataScience #Python #AIEvaluation #Benchmarking #ResearchJobs #RemoteJobs #USAJobs #Masters #PhD
