← Back to results

inference-optimization jobs in San Diego

$140,800 – $211,200 · Posted 6 days ago

Design, develop, and optimize machine learning systems for production AI platforms, focusing on model development, inference optimization, and scalable ML infrastructure. Responsibilities include building ML pipelines and frameworks, optimizing model inference for latency and cost, integrating LLMs into APIs and microservices, and designing data pipelines for preprocessing and feature engineering.… The role requires strong software engineering fundamentals combined with deep ML expertise, working across PyTorch/TensorFlow, model-serving systems, distributed computing, and GPU environments.

San DiegoLast seen 4 days ago
$122,800 – $184,200 · Posted 6 days ago

Design, develop, and optimize production machine learning systems including model development, inference optimization, and scalable ML infrastructure. Build end-to-end ML pipelines from training through deployment, optimize model inference for latency and cost, integrate LLM/ML models into APIs and microservices, and engineer data pipelines for preprocessing and validation.… Requires strong software engineering fundamentals with deep ML expertise, proficiency in Python and systems languages, and experience with ML frameworks, model serving, and distributed computing.

San DiegoLast seen 4 days ago
$137,100 – $227,000 · Posted 7 days ago

Lead end-to-end AI platform architecture and infrastructure at enterprise scale, designing and implementing generative AI capabilities including LLMs, RAG systems, and agentic AI with a focus on model serving, inference optimization, and production deployment. Establish MLOps/LLMOps best practices, build automated CI/CD pipelines tailored for AI/ML applications, and architect vector database integration and cloud AI/ML services.… Mentor engineering teams, drive cross-functional collaboration between software, product, and data teams, and continuously evaluate and optimize AI platform performance, scalability, and reliability.

San DiegoLast seen 5 days ago
Posted 9 days ago

AI Research Scientist role focused on developing and optimizing large language models for on-device deployment on smartphones, IoT, and automotive systems. Responsibilities include researching and prototyping model compression, quantization, pruning, distillation, and inference optimization techniques.… The role involves collaboration with silicon, software, and product teams to design experiments, benchmark models, and translate research into production-ready solutions for 5G and edge AI applications.

San DiegoLast seen 7 days ago
$167,200 – $229,900 · Posted 12 days ago

Senior Machine Learning Engineer to design and ship production voice and conversational AI agents within AppFolio's Realm-X platform. You will architect real-time, multi-turn agent pipelines that balance reasoning depth against latency, lead a small pod of ML and platform engineers, and define quality metrics and evaluation harnesses.… Required expertise includes shipped production experience with agent frameworks (LangChain, LangGraph), voice stacks (STT/TTS/Voice-to-Voice models), LLM reasoning and tool use, Twilio, AWS, expert Python with async and WebSocket streaming, and demonstrated team leadership.

San DiegoLast seen 10 days ago
$150,000 – $220,000 · Posted 12 days ago

Build applied AI systems for Crucible, Firestorm's manufacturing operations software. You'll productionize ML and optimization models, integrate foundation models and LLMs into workflows, and design end-to-end AI capabilities—from model selection and evaluation through reliable production deployment across cloud, air-gapped, and edge environments.… This is a hands-on engineering role requiring 5+ years shipping production ML/AI systems, strong Python and software engineering fundamentals, and deep experience with LLMs, transformer models, and production ML monitoring.

San DiegoLast seen 10 days ago
$99,500 – $149,300 · Posted 14 days ago

Develop, optimize, and validate AI/ML solutions that leverage Qualcomm's AI inference accelerators for data center and hybrid applications. Design and implement GenAI and LLM applications, perform model benchmarking, optimize deployment strategies, and drive system-level architecture decisions.… Apply deep expertise in ML frameworks, system performance profiling, and parallel computing to ensure best-in-class inference performance, power efficiency, and scalability across heterogeneous hardware.

San DiegoLast seen 12 days ago
$162,000 – $243,000 · Posted 19 days ago

As an AI Performance Engineer at Qualcomm, you will create and implement machine learning techniques, frameworks, and tools for efficient discovery and deployment of ML solutions across mobile, edge, auto, and IoT products. You will model, architect, and develop advanced ML hardware co-designed with software, optimize software for AI model deployment on hardware (kernels, compilers, model efficiency tools), and develop ML techniques into products.… You will conduct experiments to train and evaluate ML models, work independently with minimal supervision, and provide technical guidance to team members. The role requires 4+ years of hardware/software/systems engineering experience, proficiency with ML frameworks (TensorFlow, PyTorch, Keras), embedded systems optimization, and programming languages suited for ML (Python, C++, R).

San DiegoLast seen 17 days ago
$200,000 – $250,000 · Posted 20 days ago

AppFolio is hiring a Staff Machine Learning Engineer to design, build, and operate their ML platform on AWS, supporting training, fine-tuning, inference, RAG, and cost optimization across the organization's AI initiatives. You'll partner with applied AI and research teams to productionize prototypes, maintain multi-provider LLM reliability (OpenAI, Google, Anthropic), and operate AI safety guardrails and authorization layers.… The role requires production-scale ML infrastructure experience on AWS (ECS, SageMaker, GPU fleets), deep knowledge of model serving and inference optimization, hands-on language model training, and demonstrated cost discipline across AI workloads.

San DiegoLast seen 19 days ago
$158,400 – $237,600 · Posted 1 month ago

Lead end-to-end model optimization and transformation for large language models, vision language models, diffusion, and multimodal models on Qualcomm inference accelerators. Architect PyTorch-based optimization strategies, drive graph capture and deployment using PyTorch/ONNX/torch.compile, and design fusion kernels using Triton or similar DSLs.… Partner with compiler, performance, and accuracy teams to co-design lowering strategies, optimize transformer-specific patterns (KVcache, decoding, long context), and scale distributed inference across multi-core and multi-device systems. Requires expert-level PyTorch proficiency, deep transformer architecture knowledge, hands-on experience with torch.compile/TorchDynamo, and strong foundation in ML accelerators and distributed systems.

San DiegoLast seen 28 days ago
$137,100 – $227,000 · Posted 1 month ago

Lead end-to-end AI platform architecture and design for generative AI solutions, including LLMs, RAG systems, and agentic AI. Build MLOps/LLMOps pipelines with automated CI/CD, model serving infrastructure, and vector database integration on enterprise cloud platforms.… Establish best practices, mentor engineering teams, and drive seamless AI application integration into production systems while optimizing for scalability, security, and performance.

San DiegoLast seen 1 month ago
$115,300 – $160,100 · Posted 1 month ago

Senior AI Software Engineer responsible for designing, developing, and deploying LLM-powered applications and AI-assisted development workflows within a large-scale enterprise Linux environment (~6M+ LOC, primarily C/C++). Build high-performance GPU-based inference pipelines using vLLM and modern frameworks, develop agentic AI workflows with RAG and tool calling, and integrate LLMs with vector databases and enterprise systems into production.… Collaborate across software engineers, AI researchers, and platform teams to productionize AI services while optimizing for performance, latency, scalability, and operational efficiency. Requires 4+ years software development, strong Linux/Red Hat experience, advanced C/C++, Python, LLM/RAG expertise, and GitLab CI/CD workflows.

San DiegoLast seen 17 days ago
$158,400 – $237,600 · Posted 1 month ago

Staff/Sr. Staff Software Engineer role focused on AI inference optimization on Snapdragon platforms, including model optimization, quantization, graph transformations, and runtime execution for LLMs, LVMs, and LMMs.… You will design and implement graph lowering and optimization techniques within ONNX Runtime, ExecuTorch, and Qualcomm AI Stack SDK, working across ML algorithms, inference systems, and hardware integration. The role requires 6–8+ years of software development experience, 3+ years in AI/ML inference or model optimization, deep expertise in Python and C/C++, PyTorch/ONNX, and transformer architectures. You will mentor junior engineers, drive features end-to-end, and collaborate across ML Research, hardware, product, and QA teams.

San DiegoLast seen 18 days ago
$140,800 – $211,200 · Posted 1 month ago

Qualcomm AI Research seeks a Senior AI Research Quantization Engineer to develop algorithms for efficient generative AI, LLMs, and multimodal models optimized for on-device deployment. The role focuses on advanced quantization techniques, model compression, inference optimization (batching, KV caching, speculative decoding), and system prototyping using Python and PyTorch.… You will collaborate across hardware, software, and systems teams to enable state-of-the-art models to run on power- and memory-constrained devices including smartphones, autonomous vehicles, robotics, and IoT platforms. A Bachelor's degree plus 2+ years of related engineering experience (or Master's with 1+ year, or PhD) is required.

San DiegoLast seen 20 days ago
$122,800 – $184,200 · Posted 1 month ago

Develop software for the Qualcomm AI Stack SDKs (QAIRT and Genie) to enable execution of generative AI models and large language models on Snapdragon platforms. You will validate, optimize, and debug machine learning inference solutions, collaborating with cross-functional teams on feature development, performance analysis, and system reliability.… The role requires expertise in machine learning frameworks, embedded systems optimization, programming languages like Python or C++, and low-level OS/hardware interactions.

San DiegoLast seen 1 month ago
$158,400 – $237,600 · Posted 1 month ago

Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions.… Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.

San DiegoLast seen 2 days ago
$140,800 – $211,200 · Posted 1 month ago

As a Senior Software Engineer focused on AI Tools, you will reauthor and optimize generative AI models (LLMs like Llama, Phi, Qwen, and multimodal models) for efficient execution on Qualcomm's on-device hardware. You'll translate hardware constraints into model-level transformations that preserve accuracy while enabling edge deployment, integrate inference acceleration techniques, and collaborate with compiler and quantization teams to move research prototypes into production.… The role requires deep implementation-level knowledge of generative AI architectures, strong Python proficiency in large typed codebases, and hands-on experience optimizing inference for resource-constrained environments.

San DiegoLast seen 24 days ago
$178,400 – $267,600 · Posted 1 month ago

Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton.… You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.

San DiegoLast seen 5 days ago
$141,200 – $278,300 · Posted 1 month ago

Lead the design and delivery of enterprise AI platforms and applications on Google Cloud, leveraging Vertex AI, Gemini, and cloud-native technologies. Design, fine-tune, and govern LLM solutions; build RAG and agentic systems; and define end-to-end architectures spanning data pipelines, feature engineering, model lifecycle, APIs, and MLOps/LLMOps.… Architect cloud-native applications on GKE, Cloud Run, and managed services while implementing security, governance, and production-grade monitoring for AI/ML systems at scale.

San DiegoLast seen 1 month ago
$179,200 – $268,800 · Posted 1 month ago

As a Staff Product Data Manager for AI Platforms & Systems Solutions at Qualcomm, you will own the product strategy, definition, and lifecycle for AI platform products and specifications spanning model enablement, agentic systems, orchestration runtimes, and open protocols. You will develop comprehensive plans of record including schedules, budgets, and resources, directing products from conception through end of life while functioning as a central resource across systems, engineering, standards, quality, marketing, and partner teams.… The role requires deep fluency in the Windows edge-AI stack, with technical depth across AI/ML models, agentic systems, orchestration engines, standards development, and hands-on experience delivering Windows-based AI products leveraging local accelerators (CPU, GPU, NPU). You will translate emerging model, agentic, and orchestration trends into differentiated product and specification roadmaps.

San DiegoLast seen 1 month ago
$111,300 – $166,900 · Posted 1 month ago

As a Senior Systems Engineer for Data Center AI at Qualcomm, you will research, develop, optimize, and validate AI/ML solutions that integrate Qualcomm's hardware accelerators with software and ecosystem to deliver high-performance inference in datacenter environments. You will design and deploy Gen AI and LLM applications, implement fine-tuning and distillation techniques, perform model benchmarking, and drive system-level architecture decisions.… You will collaborate across functional teams to meet system-level requirements while staying current with advances in AI/ML models and hardware. Required skills include strong Python proficiency, expertise in ML frameworks (PyTorch, TensorFlow), deep understanding of ML deployment and system performance profiling, and experience with parallel computing.

San DiegoLast seen 8 days ago