Role Summary
Microsoft's AI products serve hundreds of millions of users across languages and markets. We need people who can define what a great AI response looks like in German, not just label whether one is correct.
As a Language Evaluation Specialist, you will be the quality authority for German. You'll start by evaluating AI outputs hands-on, then grow into shaping the evaluation standards themselves — working directly with product and engineering teams to close the gap between what our AI produces and what local users actually need.
Key Responsibilities
- Define quality for your language market. Evaluate AI-generated content against real user needs — not just surface-level correctness, but whether the response truly serves someone working in that language and culture.
- Turn observations into standards. Identify patterns across evaluations, extract reusable principles, and collaborate with PMs to build evaluation guidelines that scale beyond your own judgment.
- Evaluate cases with local user needs at the center. Score, review, and quality-check AI-generated cases in your target language — grounded in the linguistic nuances, cultural context, and real user expectations of your market.
- Collaborate with LLMs to drive quality at scale. Work alongside AI-assisted evaluation pipelines — reviewing model outputs, identifying where automated judgments fall short, and helping build feedback loops that continuously improve both human and model evaluation quality.
Required Qualifications
- Native-level fluency in German (reading and writing), with deep understanding of local culture, communication norms, and user expectations.
- Independent judgment with strong logical and analytical thinking: able to form your own view of what makes an AI response good or bad, articulate why, and abstract from concrete cases to generalizable evaluation criteria.
- Self-motivated with a quality-first mindset. Able to work independently, follow established guidelines with discipline, and stay engaged with detail-heavy work. Passion for AI products and how they can be improved is a strong plus.
- Professional working proficiency in English (written and verbal), able to collaborate effectively in a global, cross-timezone team environment.
Preferred Qualifications
- Experience in linguistic QA, localization QA, search relevance, or human evaluation programs — especially roles involving evaluation criteria design or refinement, not just execution.
- Hands-on familiarity with AI/LLM products (ChatGPT, Copilot, Gemini, etc.) as a regular user, with a sense of where they perform well and where they fall short.
- Understanding of how professionals in the target market work — their tools, workflows, and communication habits — so you can evaluate AI outputs from the perspective of the people who actually use them.
- Experience working in cross-cultural, cross-timezone team environments.
Job Types: Full-time, Contract
Contract length: 12 months
Pay: 2.500,00€ - 3.000,00€ per month
Language:
Work Location: Remote