Irvine, CA
Agentic AI/ML Engineer, Multimodal
What the posting asks for
- Names
- PyTorch, Python, vLLM
- Doctorate
- Mentioned, without saying whether it is needed
- Experience
- 2+ years asked for
Read out of the employer's own description. Absence means the posting does not say, not that the answer is no.
The role
Who are We? Field AI is transforming how robots interact with the real world. We are building risk-aware, reliable, and field-ready AI systems that address the most complex challenges in robotics, unlocking the full potential of embodied intelligence. We go beyond typical data-driven approaches or pure transformer-based architectures, and are charting a new course, with already-globally-deployed solutions delivering real-world results and rapidly improving models through real-field applications. Learn more at https://fieldai.com. About the Job Our Field Foundation Model (FFM) powers a global fleet of autonomous robots that capture massive streams of multimodal data across diverse, dynamic environments every day. As part of the Insight Team our mission is to transform this raw, multimodal data into actionable insights that empower our customers and engineers to deliver value. Field-insight Foundation Model (FiFM) is at the core of how we transform multimodal data from autonomous robots into actionable insights. As an AI/ML Engineer on the FiFM team, you will drive research and model development for one of Field AI’s most ambitious initiatives. Your work will span computer vision, vision-language models (VLMs), multimodal scene understanding, and long-memory video analysis and search, with a strong emphasis on agentic AI (tool use, memory, multimodal retrieval-augmented generation).This is a full-cycle ML role: you’ll curate datasets, fine-tune and evaluate models, optimize inference, and deploy them into production. It’s a blend of applied research and engineering, requiring creativity, rapid experimentation, and rigorous problem-solving. While FiFM is your primary focus, you’ll also contribute to broader perception and insight-generation initiatives across Field AI.
What You’ll Get To Do:
- Train and fine-tune million- to billion-parameter multimodal models, with a focus on computer vision, video understanding, and vision-language integration.
- Track state-of-the-art research, adapt novel algorithms, and integrate them into FiFM.
- Curate datasets and develop tools to improve model interpretability.
- Build scalable evaluation pipelines for vision and multimodal models.
- Contribute to model observability, drift detection, and error classification.
- Fine-tune and optimize open-source VLMs and multimodal embedding models for efficiency and robustness.
- Build and optimize Multi-VectorRAG pipelines with vector DBs and knowledge graphs.
- Create embedding-based memory and retrieval chains with token-efficient chunking strategies.
What You Have:
- Master’s/Ph.D. in Computer Science, AI/ML, Robotics, or equivalent industry experience.
- 2+ years of industry experience or relevant publications in CV/ML/AI.
- Strong expertise in computer vision, video understanding, temporal modeling, and VLMs.
- Proficiency in Python and PyTorch with production-level coding skills.
- Experience building pipelines for large-scale video/image datasets.
- Familiarity with AWS or other cloud platforms for ML training and deployment.
- Understanding of MLOps best practices (CI/CD, experiment tracking).
- Hands-on experience fine-tuning open-source multimodal models using HuggingFace, DeepSpeed, vLLM, FSDP, LoRA/QLoRA.
- Knowledge of precision tradeoffs (FP16, bfloat16, quantization) and multi-GPU optimization.
- Ability to design scalable evaluation pipelines for vision/VLMs and agent performance.
The Extras That Set You Apart:
- Experience with Agentic/RAG pipelines and knowledge graphs (LangChain, LangGraph, LlamaIndex, OpenSearch, FAISS, Pinecone).
- Familiarity with agent operations logging and evaluation frameworks.
- Background in optimization: token cost reduction, chunking strategies, reranking, and retrieval latency tuning.
- Experience deploying models under quantized (int4/int8) and distributed multi-GPU inference.
- Exposure to open-vocabulary detection, zero/few-shot learning, multimodal RAG.
- Knowledge of temporal-spatial modeling (event/scene graphs).
- Experience deploying AI in edge or resource-constrained environments.
Published by FieldAI on their own careers page and reproduced here unedited. Read it at FieldAI.
Apply at FieldAI → Applications go directly to FieldAI. This board does not sit in between, take a fee from you, or see your application.
What this listing does not tell you
Listed 378 days, which is longer than most. Of the 114 evals and benchmarks roles this board has watched from listing to removal, 99% were gone from their employer's careers page by day 378, and the median came down after 56 days. That is a description of other listings that have already ended, not a prediction about this one: this board records when a listing disappears, never why, and a posting still up is not on a clock it can see.
Pay is not confirmed for this role. Check the employer’s posting for current compensation. Explore published bands from other roles →
FieldAI has 5 roles open on this board, 3 of them in evals and benchmarks.
Free. One email on Thursdays.
More roles like this
Matched by discipline, title, listed location and work arrangement.
Same discipline: Evals & Benchmarks · Shared listed location
Tokyo
Same discipline: Evals & Benchmarks
Sunnyvale, CA
Same discipline: Evals & Benchmarks
Cambridge, MA
Same discipline: Evals & Benchmarks
Paris · London
Same discipline: Evals & Benchmarks