AI Systems Research and Development Engineer - LLM Inference Systems & Optimization

We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in **LLM inference systems and optimization**. Our mission is to build the next generation of **high-performance and intelligent inference systems**. We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads. Our work spans the full inference stack-from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as **adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization** to push the frontier of latency, throughput, scalability, and cost. Beyond optimizing individual models, we are building **intelligent and adaptive inference systems** that can automate performance optimization-rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace **AI-native engineering**, using AI not only as the workload we optimize, but also as a tool to accelerate system development, experimentation, debugging, optimization, and adaptation to new models. Our goal is to accelerate both the **speed of inference and the agility of inference development**. Recent innovations from Snowflake AI Research include **Arctic Inference**, our open-source inference system, and technologies such as **Shift Parallelism**, which dynamically adapts parallelism to workload characteristics; **SwiftKV**, which reduces redundant prefill computation; **Arctic Speculator and SuffixDecoding** for fast speculative decoding; **Jacobi Forcing** for causal parallel decoding; and **Semi-Persistence** for fast model swapping and dynamic multi-model serving. This is an exciting opportunity to collaborate with a world-class team, including founding members of DeepSpeed, vLLM, and TensorFlow. Together, we will push the boundaries of AI systems and bring cutting-edge research into production-scale AI. **Responsibilities** - Design and develop **high-performance LLM inference systems**, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels. - Develop novel techniques to improve **inference latency, generation speed, throughput, memory efficiency, scalability, and cost**. - Explore advanced inference techniques including **speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization**. - Develop **adaptive and intelligent inference systems** that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments. - Apply **AI-driven and AI-native approaches to systems engineering**, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning. - Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production. - Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism. - Develop efficient approaches for **multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling**. - Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components. - Explore **model-system co-design**, including model or post-training techniques that unlock substantially more efficient inference. - Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution. - Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production. - Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences. **Requirements** - Bachelor's degree in Computer Science, Electrical Engineering, or a related field. A Master's degree or PhD is preferred. - 5+ years of experience in one or more of the following areas: **LLM inference systems, distributed AI systems, GPU systems, or high-performance computing**. - Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models. - Hands-on experience with modern **LLM inference and serving frameworks**, such as **vLLM, SGLang, TensorRT-LLM**, or similar systems. - Experience designing, extending, or optimizing inference runtimes, including areas such as **scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving**. - Strong understanding of GPU architectures and experience with **CUDA, Triton**, or similar GPU programming environments. - Experience with performance-oriented libraries and frameworks such as **CUTLASS, cuBLAS, cuDNN**, or related technologies. - Experience profiling and diagnosing end-to-end system performance using **Nsight Systems, Nsight Compute**, or equivalent tools. - Demonstrated ability to operate as an **independent problem identifier and solver**-recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity. - Strong ability to work across **model, runtime, distributed system, and hardware layers** and reason about end-to-end performance tradeoffs. - Experience using **AI-native engineering approaches** to accelerate software development, experimentation, debugging, optimization, or system adaptation is a strong plus. - Excellent communication skills and the ability to collaborate effectively across research, engineering, and product teams. Snowflake is growing fast, and we're scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake. How do you want to make your impact? For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

Back to blog

Common Interview Questions And Answers

1. HOW DO YOU PLAN YOUR DAY?

This is what this question poses: When do you focus and start working seriously? What are the hours you work optimally? Are you a night owl? A morning bird? Remote teams can be made up of people working on different shifts and around the world, so you won't necessarily be stuck in the 9-5 schedule if it's not for you...

2. HOW DO YOU USE THE DIFFERENT COMMUNICATION TOOLS IN DIFFERENT SITUATIONS?

When you're working on a remote team, there's no way to chat in the hallway between meetings or catch up on the latest project during an office carpool. Therefore, virtual communication will be absolutely essential to get your work done...

3. WHAT IS "WORKING REMOTE" REALLY FOR YOU?

Many people want to work remotely because of the flexibility it allows. You can work anywhere and at any time of the day...

4. WHAT DO YOU NEED IN YOUR PHYSICAL WORKSPACE TO SUCCEED IN YOUR WORK?

With this question, companies are looking to see what equipment they may need to provide you with and to verify how aware you are of what remote working could mean for you physically and logistically...

5. HOW DO YOU PROCESS INFORMATION?

Several years ago, I was working in a team to plan a big event. My supervisor made us all work as a team before the big day. One of our activities has been to find out how each of us processes information...

6. HOW DO YOU MANAGE THE CALENDAR AND THE PROGRAM? WHICH APPLICATIONS / SYSTEM DO YOU USE?

Or you may receive even more specific questions, such as: What's on your calendar? Do you plan blocks of time to do certain types of work? Do you have an open calendar that everyone can see?...

7. HOW DO YOU ORGANIZE FILES, LINKS, AND TABS ON YOUR COMPUTER?

Just like your schedule, how you track files and other information is very important. After all, everything is digital!...

8. HOW TO PRIORITIZE WORK?

The day I watched Marie Forleo's film separating the important from the urgent, my life changed. Not all remote jobs start fast, but most of them are...

9. HOW DO YOU PREPARE FOR A MEETING AND PREPARE A MEETING? WHAT DO YOU SEE HAPPENING DURING THE MEETING?

Just as communication is essential when working remotely, so is organization. Because you won't have those opportunities in the elevator or a casual conversation in the lunchroom, you should take advantage of the little time you have in a video or phone conference...

10. HOW DO YOU USE TECHNOLOGY ON A DAILY BASIS, IN YOUR WORK AND FOR YOUR PLEASURE?

This is a great question because it shows your comfort level with technology, which is very important for a remote worker because you will be working with technology over time...

Other Jobs To Apply

Entry Level Graphic Designer

Flight Attendant (PT + FT Available)

Teen-Friendly Remote Online Data Entry Specialist – Flexible Hours, Competitive Pay, Earn From Home with arenaflex

Amazon Picker/Packer

Apple Mac Management JAMF Engineer

Project Administrator- (Corporate Coding Resources) REMOTE

Remote Entry-Level Typing Specialist, $15/hr, No Experience Required – Full/Part-Time

Fulfillment Center Area Manager

Hiring Now - Work from Home - No Experience

Work from Home Product Testing $25-$45 per hour

Delivery Driver - Flexible Shift

Onsite Medical & Safety Specialist

Independent Delivery Driver - Flexible Scheduling

Area Manager, Early Career (2027) – GA, KY, TN (Fulfillment & Operations)

Cloud Data Center Technician - 24/7 Ops & Travel

Dock Associate in Tracy, California ($20.00)

Mixing Center - Warehouse Associate

flex driver

Amazon DOT Delivery Driver(Professional Driving Experience Required) $22.25

Delivery Driver Helper (Amazon Packages)

Area Operations Manager - Fast-Paced Fulfillment

Medical Assistant - Hybrid (Work From Home/Clinic)

Customer Service Rep - Work From Home (PT/FT)

[Remote-Position] YouTube Content Moderator Jobs Remote $27Hr

Claims Specialist- Anchorage, AK (Hybrid)

Customer Service Agent - Graveyard Shift (Work From Home)

Customer Service - Work from home ($250 Bonus)

Weekend driver/courier (part-time) | North Miami Beach

Easy Online Typing Jobs - Work at Your Own Schedule

Immediate Start Costco Working From Home , Careers At Costco

Service Desk Analyst I — Onsite IT Support & Onboarding

Join Today: Delta Airlines Careers Remote Customer Service, Delta

Investment Finance Expert - Valuation - AI Trainer

Delta Airlines Online Careers Remote Jobs At Home (No Experience/Entry Level)

Remote Entry-Level Customer Support – No Experience Required

Netflix Remote Jobs No Experience $30/Hour -

Entry Level Chat Support Specialist (Fully Remote)

Remote Out of Office Position / Data Entry (Hiring Immediately)

Part-time Evening Phone Coordinator

Customer Service Representative Agent Work From Home - Part Time Focus Group Panelists

_Fully Remote Position (Flexible & Beginner Friendly) Start ASAP + Bonuses

Delta Remote Jobs (Data Entry)- Hiring Now

Programmer - AI Trainer

AMAZON Walker Delivery

Amazon Customer Service - United States - Work From Home

Home office - Nmet nyelv? banking support munkatrs (m/w/d)

Fulfillment Center Lead - Operations & Training

Frito-Lay - Warehouser/Material Handler $18-$26/hr

Distribution Analyst

Hospital Response Advocate- On Call