← Back to results

vllm jobs in San Diego

$77,600 – $176,000 · Posted 6 days ago

As an ML engineer at Booz Allen Hamilton, you'll design, build, and deploy production-grade AI/ML systems for Defense and Intelligence clients. You'll work across the full ML lifecycle—from model development in deep learning, computer vision, NLP, and signal processing through MLOps, containerization, and cloud deployment.… The role requires 2+ years of ML engineering or data science experience, proficiency with frameworks like TensorFlow and PyTorch, and the ability to obtain a Secret clearance.

San DiegoLast seen 5 days ago
$140,800 – $211,200 · Posted 7 days ago

Design, develop, and optimize machine learning systems for production AI platforms, focusing on model development, inference optimization, and scalable ML infrastructure. Responsibilities include building ML pipelines and frameworks, optimizing model inference for latency and cost, integrating LLMs into APIs and microservices, and designing data pipelines for preprocessing and feature engineering.… The role requires strong software engineering fundamentals combined with deep ML expertise, working across PyTorch/TensorFlow, model-serving systems, distributed computing, and GPU environments.

San DiegoLast seen 5 days ago
$122,800 – $184,200 · Posted 7 days ago

Design, develop, and optimize production machine learning systems including model development, inference optimization, and scalable ML infrastructure. Build end-to-end ML pipelines from training through deployment, optimize model inference for latency and cost, integrate LLM/ML models into APIs and microservices, and engineer data pipelines for preprocessing and validation.… Requires strong software engineering fundamentals with deep ML expertise, proficiency in Python and systems languages, and experience with ML frameworks, model serving, and distributed computing.

San DiegoLast seen 5 days ago
$223,600 – $335,400 · Posted 14 days ago

Principal Software Engineer role on Qualcomm's Cloud AI team focused on LLM serving and inference acceleration. The engineer will design, optimize, and deploy high-performance software across the product lifecycle, from R&D through commercial deployment, with expertise in serving frameworks (vLLM), PyTorch, neural network optimization, and multi-core/SoC architecture performance modeling.… Strong C++/Python development skills, deep LLM/multi-modal model understanding, and experience with machine learning accelerators are required; experience with compiler technology, performance analysis, and commercial software delivery at scale is essential.

San DiegoLast seen 12 days ago
$200,000 – $250,000 · Posted 21 days ago

AppFolio is hiring a Staff Machine Learning Engineer to design, build, and operate their ML platform on AWS, supporting training, fine-tuning, inference, RAG, and cost optimization across the organization's AI initiatives. You'll partner with applied AI and research teams to productionize prototypes, maintain multi-provider LLM reliability (OpenAI, Google, Anthropic), and operate AI safety guardrails and authorization layers.… The role requires production-scale ML infrastructure experience on AWS (ECS, SageMaker, GPU fleets), deep knowledge of model serving and inference optimization, hands-on language model training, and demonstrated cost discipline across AI workloads.

San DiegoLast seen 20 days ago
$77,600 – $176,000 · Posted 22 days ago

Design, create, and implement production-grade AI/ML systems for Defense and Intelligence clients. You'll architect scalable machine learning solutions handling fast-moving data, deploy models across multiple data modalities using frameworks like TensorFlow and PyTorch, and integrate solutions into mission environments.… The role requires 2+ years of ML engineering or data science experience, proficiency with ML frameworks and programming, and hands-on work in deep learning, computer vision, NLP, or signal processing.

San DiegoLast seen 20 days ago
$130,000 – $160,000 · Posted 24 days ago

Mid-level full-stack software engineer responsible for developing, deploying, and troubleshooting distributed applications across application, OS, networking, and hardware layers. Must have 5+ years of professional experience with Python, modern front-end frameworks, RESTful APIs, Linux (RHEL preferred), containerization (Podman/OCI), Ansible automation, and Git.… Role emphasizes independent ownership of features from requirements through testing, strong debugging across system and software components, and integration with external systems and hardware interfaces.

San DiegoLast seen 22 days ago
$130,000 – $160,000 · Posted 26 days ago

Mid-level full-stack software engineer responsible for developing, deploying, and troubleshooting distributed applications across application software, operating systems, networking, virtualization, and hardware interfaces. The role requires 5+ years of professional software development experience with proficiency in Python, modern front-end frameworks, RESTful APIs, Linux administration, containerization (Podman/OCI), Git, and configuration automation (Ansible).… Strong emphasis on independent ownership of features from requirements through implementation and testing, with preferred experience in LLMs, generative AI, agentic workflows, and AI-tool integration for engineering tasks.

San DiegoLast seen 24 days ago
$200,000 – $250,000 · Posted 1 month ago

Lead machine learning strategy and development for AppFolio's Leasing products, owning the ML roadmap and autonomous leasing agent architecture. Build evaluation frameworks, model quality infrastructure, and establish ML standards across the Leasing Engineering team while ensuring production-grade reliability, SLOs, and observability.… Translate research into shipped features by evaluating fine-tuning approaches, RAG patterns, and agentic systems; operate with production discipline on a SaaS platform serving real customer workflows.

San DiegoLast seen 15 days ago
$198,500 – $297,700 · Posted 1 month ago

Lead the design, development, and operation of Qualcomm's enterprise AI platform, supporting agentic AI, model lifecycle management, and multi-cloud ML/inference infrastructure. Own core platform services including identity/RBAC, secrets, service meshes, observability, vector stores, and model gateways across on-prem GPU clusters and managed cloud services.… Manage a ~10-engineer global team (platform, SRE, MLOps/LLMOps), drive incident response and continuous improvement, and partner with product and security teams on AI governance and high-impact use cases.

San DiegoLast seen 17 days ago
$115,300 – $160,100 · Posted 1 month ago

Senior AI Software Engineer responsible for designing, developing, and deploying LLM-powered applications and AI-assisted development workflows within a large-scale enterprise Linux environment (~6M+ LOC, primarily C/C++). Build high-performance GPU-based inference pipelines using vLLM and modern frameworks, develop agentic AI workflows with RAG and tool calling, and integrate LLMs with vector databases and enterprise systems into production.… Collaborate across software engineers, AI researchers, and platform teams to productionize AI services while optimizing for performance, latency, scalability, and operational efficiency. Requires 4+ years software development, strong Linux/Red Hat experience, advanced C/C++, Python, LLM/RAG expertise, and GitLab CI/CD workflows.

San DiegoLast seen 18 days ago
$158,400 – $237,600 · Posted 1 month ago

Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions.… Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.

San DiegoLast seen 2 days ago
$140,800 – $211,200 · Posted 1 month ago

Design, develop, and optimize machine learning systems and models for production AI platforms, with focus on inference optimization, scalable ML infrastructure, and deployment. Responsibilities include building ML pipelines, optimizing model inference across hardware environments, integrating LLMs and models into APIs and microservices, designing data pipelines for ingestion and feature engineering, and collaborating cross-functionally on end-to-end ML solutions.… Requires strong software engineering fundamentals combined with deep ML expertise, proficiency in Python and at least one systems language (C++, Rust, or Go), and solid understanding of ML frameworks, transformer architectures, and model deployment systems.

San DiegoLast seen 1 month ago
$140,800 – $211,200 · Posted 1 month ago

Design, develop, and optimize machine learning systems for production AI platforms, including model development, inference optimization, and scalable ML infrastructure. Build training-to-deployment pipelines, optimize model serving for latency and cost, and integrate LLMs and generative AI models into microservices and APIs.… Develop data pipelines for ingestion, preprocessing, and feature engineering while collaborating cross-functionally with product, platform, and hardware teams to deliver end-to-end ML solutions.

San DiegoLast seen 1 month ago
Posted 1 month ago

Design, build, and ship LLM-powered capabilities end to end—from prototyping and fine-tuning models to deploying production agents and retrieval systems. Own the full stack: prompt and context engineering, multi-step agent design with tool calling, RAG systems (embeddings, chunking, hybrid search, reranking), fine-tuning on multi-GPU with LoRA/QLoRA, evaluation and tracing infrastructure, and clean APIs for other engineers.… Work in a secure, distributed environment where the platform runs on customer compute, cloud, or hybrid setups, requiring expertise with both commercial and self-hosted models.

San DiegoLast seen 12 days ago
$94,200 – $141,200 · Posted 1 month ago

As a Product Software Engineer, you will develop and integrate software for Qualcomm products spanning smartphones, computing devices, automotive infotainment, and IoT systems. You will own software integration tasks including component integration, build automation, version control management, and baseline releases.… You will perform sanity testing on-target and in simulation environments, debug failures using tools like JTAG and ADB, automate test scenarios, and collaborate with cross-functional engineering teams on design and code reviews.

San DiegoLast seen 1 month ago