LLM Red Team & Benchmark Evaluation Specialist

24-MAG·New York, New York·Posted 2h ago·via Talent.com

RegionUSA
Apply Now

Job description

Probe frontier AI models across coding, machine learning, analysis, and multi-step agentic tasks.Identify subtle errors, vulnerabilities, edge cases, and misleadingly plausible outputs.Investigate situations where models appear capable while reaching incorrect or unsupported conclusions.Design repro...

first seen 2026-07-27 16:48:02 · last verified 2026-07-27 16:48:02


pentestcareers.com // breach the job market