Senior Machine Learning Engineer, Proactive
Apple
Santa Clara, CA · Onsite · Full Time
Posted
Job description
At Apple, machine learning powers experiences that anticipate what people need before they ask. We're looking for a Machine Learning Engineer to help build the next generation of intelligent search and AI experiences technology that understands user intent, context, and personal information while preserving privacy. In this role, you'll design, train, optimize, and deploy large language models, semantic retrieval systems, and ranking models that power relevant, personalized, context-aware search across Apple's ecosystem. You'll work at the intersection of search, retrieval, natural language processing, on-device AI, and generative AI to shape the future of intelligent assistants and proactive experiences.. Description You'll design, train, fine-tune, and optimize transformer-based language models for on-device deployment, and build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems that improve search quality and AI-powered experiences. You'll develop models for query understanding, intent prediction, personalization, and ranking, while researching new approaches to model compression, quantization, and low-latency inference. You'll partner with engineers, researchers, product managers, and designers to bring new AI capabilities from research into production driving technical strategy and leading projects from early exploration through large-scale deployment. This is an opportunity to explore new applications of foundation models, multimodal AI, and agentic retrieval, shaping the next generation of proactive, intelligent user experiences. Responsibilities: Design, train, fine-tune, and optimize transformer-based language models for efficient on-device deployment. Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems, along with models for query understanding, intent prediction, personalization, and ranking. Research and prototype approaches for on-device generative AI, including model compression, quantization, knowledge distillation, and low-latency inference. Analyze search relevance and user behavior to design evaluation methodologies, offline benchmarks, and online metrics that measure retrieval quality, ranking, and language model performance. Partner with engineers, researchers, product managers, and designers to bring AI capabilities from research into production, driving technical…