Staff / Senior Software Engineer (Agentic Search) - Runtime

Nebius

London · Onsite · Full Time

Posted

Job description

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Product In a rapidly evolving world, trust in AI depends on AI agents being grounded in fresh, verified real-world data. Search is the foundation that makes this possible. We are building an agent-native search platform designed specifically for AI systems rather than human users. Our product provides programmatic, low-latency, and observable search APIs that AI agents use to retrieve, filter, and reason over real-world information at scale. The Role We are looking for a Senior Software Engineer to work on the runtime systems of a novel search engine tailored for agentic AI consumption. In this role, you will focus on building low-latency, high-throughput systems that serve search queries in real time. You will work on the critical path of user-facing requests, where performance, predictability, and efficiency directly impact product quality. You will design and operate systems that handle thousands of requests per second under strict latency budgets, optimising every layer from request handling to data access and response assembly. In this position, your responsibility will be to Design, implement, and operate core runtime services for serving search queries at scale Build and optimise request flows, including query processing, retrieval orchestration, and response assembly under strict latency budgets Develop systems that maintain performance and predictability under high load Optimise CPU, memory, and data access patterns in performance-critical paths Ensure reliability, observability, and predictability across production services Build well-tested systems with clear responsibilities and interaction contracts, w…

Apply for this job