Principal Security Research Manager
Microsoft
We are building the evaluation backbone for safe, reliable, and efficient agentic engineering in Microsoft Security. This team will create systems that determine when an AI agent, model, prompt, tool, memory strategy, or orchestration pattern is ready to be used in production security and engineering workflows. The role is ideal for engineers who can operate end to end: understand the workflow, design the benchmark, build the harness, implement validators and graders, run experiments, analyze quality and cost tradeoffs, connect results to production feedback, and help teams make evidence-based release decisions.
Why this role matters: Microsoft Security is moving toward agentic engineering systems for security triage, remediation, repo readiness, and scan-to-verified-closure workflows. Evals are the trust system for that shift. They help decide whether autonomy can safely expand, whether a release should stop, and which configuration achieves the required quality, safety, reliability, latency, and cost bar with the lowest practical human-review burden.
Role mission
As a Principal Security Research Manager on the AI Evaluation Systems team, you will build the common evaluation platform and methodology used by MSec agent programs. You will work across evaluation design, platform implementation, test infrastructure, telemetry, measurement, security workflow understanding, and production learning. Your work will make agentic systems measurable, reproducible, governable, and continuously improving.
Responsibilities
- Design and build end-to-end evaluation harnesses for agentic security and engineering workflows, including triage, remediation, repo readiness, escalation, tool use, and scan-to-verified-closure paths.
- Create representative benchmark suites and golden datasets that include normal, edge, adversarial, failure-recovery, regression, and production-derived cases.
- Implement deterministic validators, automated graders, trace analyzers, result stores, comparison views, and workflow adapters that make evaluations repeatable and actionable.
- Measure task success, correctness, safety and policy compliance, failure recovery, latency, tool-call behavior, token usage, total cost per successful outcome, and human-review effort.
- Compare models, prompts, tools, memory strategies, policies, and orchestration patterns under consistent conditions and help teams understand quality-versus-efficiency tradeoffs.
- Integrate evaluations into engineering workflows, CI/CD, release gates, and decision processes so material agent changes are supported by reproducible evidence before production rollout.
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.