Research Fellowship: Agent Intelligence & Evaluation

04 Oct 2026
Apply

Voice agents fail in ways traditional software doesn't. An ASR confidence drop on a regional accent misfires a tool call, an LLM hallucinates a policy because upstream latency broke turn-taking, and support teams roll these agents back within a week without anyone able to explain what went wrong.

 What you'll work onOver 4 months, you'll take on one or two of the following, shaped by your interests.Evaluation frameworks. Text-only evals miss most of what matters in voice: barge-in, prosody, latency-induced errors, cross-turn context loss. You'll design audio-native metrics, generate adversarial conversational datasets across accents and edge cases, and build LLM-as-judge rubrics for task completion, empathy, and recovery from tool failures.End-to-end observability. Tracing a failed interaction means correlating audio packets, STT hypotheses, LLM reasoning traces, tool calls, and TTS output back to a single conversation ID. You'll help shape the schema and analysis layer that makes cascade failures visible across the stack.Self-improvement systems. Once you can measure and trace, the interesting work is closing the loop: mining production traces for failure patterns, generating targeted fine-tuning data or prompt updates, and validating that fixes hold under adversarial replay.

  • ID: #55355953
  • State: New York Newdelhi 00000 Newdelhi USA
  • City: Newdelhi
  • Salary: USD TBD TBD
  • Job type: Full-time
  • Showed: 2026-10-04
  • Deadline: 2026-12-03
  • Category: Et cetera
Apply