Skip to content
Anirudh Rao

Machine Learning Engineer · AppleBay Area, CA ·

I build ML that findsthe answer — and proves it.

Production-scale generative AI at Apple — RAG, LLM evaluation, and personalization for millions of customers. MS CS (AI) at Georgia Tech.

01Impact

Numbers from production, not notebooks.

Annual sales lift from recommendations
$80M+
LLM-as-judge evaluations processed
8M+
Files in production RAG retrieval
107K+
Fewer production issues via eval gates
40%
02Systems at AppleInternal · summarized

Three systems, millions of customers.

  1. 01

    Generative AI platform

    Production-scale RAG system for millions of customers — prompt engineering with context optimization, content-safety frameworks, and multi-model orchestration.

    • RAG
    • LLMs
    • Safety
    • Orchestration
    107K+files · 98%+ retrieval accuracy2024 — Now
  2. 02

    LLM evaluation infrastructure

    Distributed LLM-as-judge framework, automated regression pipelines over benchmark datasets, and quality gates with continuous monitoring.

    • LLM-as-judge
    • Evals
    • Regression testing
    8M+evaluations · 40% fewer prod issues2024 — Now
  3. 03

    Mac Recommendations

    Question-based recommendation engine built from the ground up, plus the Top Model system with A/B testing, Kafka analytics, and real-time buyability suppression.

    • Recommenders
    • A/B testing
    • Kafka
    $80Mannual sales lift · 95%+ accuracy2023 — 2024
03

How I think about it

Great models don’t matter until they retrieve the right context, pass the eval, and ship to real people.

head 0 · query = “retrieve” · hover a highlighted word
04Open workAll work →

Projects with the code attached.

esc
↑↓ to move · ↵ to selectk=16