← Back to results

Jobs at Xora Innovation

Posted 6 days ago

Own the complete release and deployment strategy for Elemynt's secure AI infrastructure platform, building systems that ensure every release is versioned, reproducible, observable, and safe to operate across enterprise and scientific computing environments. Design and implement deployment automation, CI/CD pipelines, artifact management, and operational visibility (logs, metrics, alerts) that work reliably in customer-managed and restricted environments. Debug incidents, identify root causes, and systematize fixes into reusable automation. Partner with product and engineering teams to embed deployment readiness into the platform from the start.

San DiegoLast seen 4 days ago
Posted 6 days ago

Design and own cloud infrastructure foundations, core platform services, and observability systems that enable reliable deployment across customer environments (on-premise, cloud, or hybrid). Build production-grade Kubernetes infrastructure, internal APIs, deployment pipelines, and monitoring/alerting layers using infrastructure-as-code. Instrument service-level objectives and health signals to ensure measurable reliability and reproducible, secure deployments across all environments.

San DiegoLast seen 4 days ago
Posted 6 days ago

Build and operate the data and ML infrastructure powering an AI platform for materials science, owning both sides: data pipelines that ingest and curate large-scale scientific output into training-ready formats, and model packaging, serving, monitoring, and CI/CD systems that move models safely from research to production across customer environments. You will design data ingestion and transformation workflows, implement validation and quality gates, package and version models with reproducible builds, run models through batch and online inference with safe rollout and rollback, monitor for drift and degradation, and build observability and internal tooling for engineering and science teams. The role requires 6+ years shipping production software with deep expertise in data systems, ML infrastructure, containers, orchestration, and observability.

San DiegoLast seen 4 days ago
Posted 6 days ago

Design, build, and ship LLM-powered capabilities end to end—from prototyping and fine-tuning models to deploying production agents and retrieval systems. Own the full stack: prompt and context engineering, multi-step agent design with tool calling, RAG systems (embeddings, chunking, hybrid search, reranking), fine-tuning on multi-GPU with LoRA/QLoRA, evaluation and tracing infrastructure, and clean APIs for other engineers. Work in a secure, distributed environment where the platform runs on customer compute, cloud, or hybrid setups, requiring expertise with both commercial and self-hosted models.

San DiegoLast seen 4 days ago
Posted 6 days ago

As Principal Software Engineer for AI & Data Platform, you will architect the data foundation for scientific and engineering R&D platforms, designing scalable data processing patterns, ML training pipelines, and intelligent workflow interfaces. You will own end-to-end responsibilities including data modeling for multi-use analytics and ML, building production training and fine-tuning pipelines, model evaluation and benchmarking, and setting engineering standards for the team. The role requires 10+ years shipping production software, expert-level Python, deep experience with large-scale data systems (object storage, analytical processing, training formats), hands-on ML pipeline development, and the ability to set technical direction in early-stage environments while implementing it yourself.

San DiegoLast seen 4 days ago
Posted 6 days ago

Build the foundational agentic AI layer for a materials-science platform, including multi-model provider abstraction, agent orchestration with stateful checkpoints, retrieval systems, prompt versioning, and comprehensive tracing and evaluation frameworks. You'll design agents that plan and reason over tool calls in production, implement human-in-the-loop safety gates, and ensure all LLM behavior remains auditable and cost-tracked across customers' secure environments. The role demands deep production experience with agentic and LLM systems: async Python, structured outputs, memory and context management, multi-step workflow orchestration, and evaluation harnesses that catch regressions before deployment.

San DiegoLast seen 4 days ago