Browse all jobs
    jupus

    Senior AI Engineer | Voice Agent Evaluation

    jupus
    Germany2 months ago
    German nice-to-have
    Data & AI
    AI Engineering
    Senior
    Remote

    Summary

    Seeking a Senior AI Engineer to own the quality and evaluation of a production voice AI system for a legal tech company's AI phone assistant. Requires strong Python, SQL, LLM prompt engineering, and experience with voice/speech systems and data evaluation. Role offers foundational impact and technical ownership in a fully remote, growth-stage environment.

    Location
    Germany
    Type
    full-time
    Level
    Senior
    Work mode
    remote

    We're seeking a driven Senior AI Engineer to join one of Europe's fastest-growing legal tech companies as we expand the core AI team behind 'Marie', our AI phone assistant.

    This role owns the quality and evaluation of a production voice AI system handling real customer calls at volume. You'll own Marie's conversation quality on top of our voice framework and providers — not the telephony stack itself. Reporting to the AI Lead, you'll work hands-on with product, engineering, and operations to make Marie measurably more reliable release over release.

    For someone who has evaluated and improved production conversational or voice systems, this is a chance to own Marie's quality loop and prove what works with data.

    Your Challenge and Opportunity in This Role

    You'll define how voice AI quality is measured and improved here — building the evaluation systems that tell us whether Marie is actually getting better, using AI-assisted review to work through real calls at scale rather than by hand. Expect strong ownership and fast iteration as you turn messy production calls into datasets, metrics, and regression tests.

    Evaluation & Quality:

    • Own Marie's conversation quality and build the evaluation loop: datasets, labels, rubrics, regression tests, and release checks

    • Use AI-assisted review to classify, label, summarize, and cluster call failures, surfacing uncertain cases for human review

    • Define and track metrics like task success, routing and capture accuracy, escalation correctness, hallucination rate, and latency

    • Run before/after comparisons for speech-recognition, TTS, model, prompt, and vendor changes to decide whether a change helped

    • Analyze transcripts, traces, and logs to separate voice-stack problems from prompt, product, and data problems

    • Improve prompts, clarification flows, fallbacks, and structured capture, then prove the impact with evals

    Cross-functional Collaboration:

    • Work with product, engineering, and operations to decide which workflows Marie should automate and which go to a human

    • Partner with engineers and vendors to debug production voice failures across telephony, speech, and latency

    • Translate complex AI and conversation-quality problems into practical decisions for technical and non-technical stakeholders

    Requirements

    Technical Expertise:

    • Experience with voice, speech, or phone systems using real audio and real users

    • Strong evaluation and data skills: defining test cases, labeling production data, building ground truth, tracking metrics, and turning failures into regression tests

    • Experience using AI tools to review messy production data at scale

    • Working understanding of speech recognition, TTS, turn-taking, interruptions, latency, and voice-agent failure modes

    • Excellent conversation-design judgment — knowing when an agent should ask, confirm, clarify, escalate, or stay quiet

    • Solid grasp of LLM behavior, prompt engineering, tool calling, structured outputs, and observability

    • Enough Python or SQL to build datasets, compare outputs, inspect traces, and automate checks

    • Experience improving AI systems in real production environments, not research settings

    Nice to Have:

    • A background in data science, ML, or applied conversational/chatbot design

    • Familiarity with voice-agent frameworks or speech providers (Pipecat, LiveKit Agents, Vapi, Retell, OpenAI Realtime, Deepgram, ElevenLabs, Cartesia) — we care more about how you evaluated them than which you used

    • Experience with eval or observability tools such as Langfuse, Braintrust, LangSmith, Promptfoo, or custom harnesses

    • Experience with legal tech, support automation, or intake, booking, and routing flows

    • Fluency or working proficiency in European languages such as German, French, Spanish, or Italian

    Note:

    • You must be located in the CET -2/+2 timezone for this role!

    Benefits

    🚀 Foundational Impact – Help build and scale the evaluation system for one of JUPUS's core AI products

    🔒 Technical Ownership – Own the quality loop behind a production AI phone assistant used by real customers

    🌱 Growth Environment – Work on production voice AI, evals, traces, and measurable product outcomes at the cutting edge of generative AI

    💼 Leadership Visibility – Flat hierarchy with direct access to the Head of Engineering and CTO, reporting to the AI Lead

    🏡 Fully Remote – Enjoy the flexibility of working from anywhere

    💻 Top Equipment – State-of-the-art MacBook plus all the accessories you need

    🏋️ Wellbeing Perks – Urban Sports Club membership to support your healthy lifestyle

    If you're passionate about building strong evals for production voice AI, using AI tools to understand messy real-world calls, and making an AI phone assistant that users rely on measurably better every day, we'd love to hear from you.

    Senior AI Engineer | Voice Agent Evaluation

    jupus · Germany

    Apply for this role

    We use analytics cookies (Umami, Vercel) and a feedback widget (Userback) to improve Joblyst. You can accept or reject non-essential cookies. Cookie policy