LLM Red Team & Benchmark Evaluation Specialist
RegionUSA
Job description
Probe frontier AI models across coding, machine learning, analysis, and multi-step agentic tasks.Identify subtle errors, vulnerabilities, edge cases, and misleadingly plausible outputs.Investigate situations where models appear capable while reaching incorrect or unsupported conclusions.Design repro...
first seen 2026-07-27 16:48:02 · last verified 2026-07-27 16:48:02
pentestcareers.com // breach the job market