Field Engineering
How customer-facing engineers at an inference and fine-tuning cloud turn technical depth into a customer outcome: discovery, sizing, benchmarks run for a customer, POCs, migrations, commercials, security reviews, and escalations.
Prerequisites: Comfort reading a latency percentile and a cost table. The technical depth (prefill and decode, KV cache, load testing math) lives in LLM Systems & Inference and Serving & Load and is linked here, not repeated.
Overview
Section titled “Overview”- For: sales engineers (SE), solutions architects (SA), and forward deployed engineers (FDE) at companies shaped like Fireworks AI, Together AI, or Baseten
- Primary references: Google SRE book, Service Level Objectives (free); The Mom Test by Rob Fitzpatrick (paid book); the vLLM benchmark CLI docs (free)
- Supplementary: The Site Reliability Workbook: Implementing SLOs (free), Postmortem Culture (free), Dev versus Delta (free, the original FDE model), The Palantirization of everything (free)
- Estimated time: 3-4 weeks at 6-8 hrs/week
Key Takeaways
Section titled “Key Takeaways”- Every call and every week of work ends in a written artifact someone else can act on. If it did not produce one, it did not happen.
- Discovery exists to produce a deployment hypothesis with numbers. A question that cannot change a number is not worth the customer’s time.
- A benchmark run for a customer is a contract in progress: their traffic shape, their region, their SLO, curves instead of single numbers, and no promise the data does not support.
- A POC exists to make a decision, so its criteria must be able to fail.
- Escalations that list what was ruled out get acted on; forwarded complaints get meetings.
How to Study
Section titled “How to Study”- Work the topics in order on one mock customer, Lexa (brief in 01). Each topic’s exercise produces one document; together they form a full engagement file.
- When a topic links into an ML track, open it and redo the math yourself. A field engineer who cannot defend a number in front of the customer’s MLE loses the room.
- Read every template aloud as if on a call. If a line would make you hesitate, you do not yet know what it asks for.
Prerequisite Graph
Section titled “Prerequisite Graph”graph LR
D[01 Discovery] --> Q[02 Qualification and Sizing]
Q --> P[03 Performance Testing Engagements]
Q --> C[06 Commercials and Security]
P --> POC[04 POC and Evaluation]
D --> POC
POC --> M[05 Migration and Cutover]
POC --> C
P --> E[07 Escalation and Handoff]
M --> E
E --> MOCK[08 Mock Engagement]
C --> MOCK
Topics
Section titled “Topics”| # | Topic | Primary Reference | Time |
|---|---|---|---|
| 01 | Discovery | The Mom Test (Fitzpatrick) + SRE book ch. 4 (free) | 3-4 days |
| 02 | Qualification and Sizing | Serving & Load (free, in repo) | 3-4 days |
| 03 | Performance Testing Engagements | vllm bench serve (free) + AIPerf (free) | 4-5 days |
| 04 | POC and Evaluation | LLM Evaluation (free, in repo) | 3 days |
| 05 | Migration and Cutover | Model Loading: tokenizers and chat templates (free, in repo) | 3 days |
| 06 | Commercials and Security | AICPA SOC 2 (free) + GDPR text (free) | 3-4 days |
| 07 | Escalation and Handoff | Postmortem Culture (free) | 2 days |
| 08 | Mock Engagement | Topics 01 to 07 applied to your own platform from the course; the role-path capstone | course pass 11 |
Role Landscape
Section titled “Role Landscape”Company facts below were checked against primary sources on October 6, 2026. Product surfaces change monthly; recheck before a call.
What the companies sell
Section titled “What the companies sell”All three sell GPU time wrapped in software. The wrapper differs: a per-token API, a dedicated deployment, a training job, or raw cluster capacity. A field engineer has to know which wrapper fits a workload and what each costs the customer in control and money.
| Surface | Fireworks AI | Together AI | Baseten |
|---|---|---|---|
| Serverless per-token API | Pay-per-token serverless models (quickstart) | Serverless inference, OpenAI-compatible (quickstart, product) | Model APIs, compatible with OpenAI Chat Completions and Anthropic Messages (beta); tool calling, structured output, JSON mode (Model APIs) |
| Dedicated / on-demand deployment | On-demand deployments on dedicated GPUs with autoscaling (on-demand quickstart, why on-demand) | Dedicated endpoints, single-tenant GPUs (docs) | Dedicated deployments of open or custom models (first model) |
| Fine-tuning | Managed: SFT, DPO, LoRA. Training API: LoRA and full-parameter, custom loops, RL (GRPO and others via SDK). RFT listed (intro) | SFT, LoRA (default) or full fine-tuning, preference tuning with DPO (overview, DPO). RL not listed in fine-tuning docs at review | Training Jobs and Loops (training, loops). Supported method list not confirmed |
| Batch / async | Batch inference (docs) | Batch API (docs) | Async inference with webhooks on any dedicated deployment, queue up to 72 h (docs). No separate batch product confirmed |
| Embeddings / rerank | Embeddings and reranking (docs) | Embeddings and rerank API; rerank models need a dedicated endpoint (rerank) | BEI engine for embeddings and reranking (BEI) |
| Custom model deploy | Custom model upload on on-demand deployments (announcement); multi-LoRA serving (blog) | Custom model upload (docs); Dedicated Container Inference for your own Docker image (docs) | Truss (package model code) and Chains (multi-step, multi-model pipelines) (Chains, Truss CLI) |
| GPU clusters | Not a headline product at review; not confirmed | H100, H200, B200, GB200 clusters with storage for training or large batch (docs, product) | Not offered as a raw cluster product; not confirmed |
| BYOC / VPC / self-hosted | Bring Your Own Cluster on customer Kubernetes, Private Preview, enterprise only, inference only (BYOC); SageMaker as a compute option (blog) | Together Enterprise Platform in customer VPC or on-prem (announcement, architecture) | Baseten Cloud, single-tenant isolated VPC, self-hosted in customer VPC, hybrid (hosting options) |
| Engine work | FireAttention custom kernels: V3 targets AMD MI300, V4 targets NVFP4 on B200 (V3, V4) | Together inference engine, Together Kernel Collection, ATLAS runtime-learning speculator (ATLAS, research) | Baseten Inference Stack: Engine-Builder-LLM on TensorRT-LLM, BIS-LLM for MoE, speculative decoding (EAGLE, MTP, n-gram), KV-aware routing (guide, BIS-LLM) |
Notes on the table:
- Vendor speedup numbers (for example FireAttention versus vLLM, ATLAS throughput) are the vendor’s own benchmarks on the vendor’s chosen workload. Treat them as hypotheses to reproduce on the customer’s traffic (03), the same rule as in LLM Serving Platforms.
- Compliance claims change and are contractual. Fireworks publishes SOC 2 Type II and HIPAA status and a no-logging default for open-model prompts (data security). Together documents its posture at privacy and security. For any customer, get the current report from the trust center, not from a blog (06).
- Together also sells Sandbox and Managed Storage (products); Baseten documents a Frontier Gateway for model labs (overview). These are outside this track.
Common to all three: an OpenAI-compatible request surface, a serverless tier for evaluation, a dedicated tier for SLOs, fine-tuning feeding deployment, and a proprietary engine layer on top of or beside vLLM, SGLang, or TensorRT-LLM. The differentiator they sell is the engine and the people. Field engineers are part of the people.
Who owns what
Section titled “Who owns what”| Role | Owns | Does not own | Typical artifact |
|---|---|---|---|
| Sales engineer (SE) | Pre-sales technical fit for one deal: demos, discovery, first sizing, security questionnaire answers, technical win | Post-sale delivery, engine internals | Discovery notes, demo, questionnaire answers, initial hypothesis |
| Solutions architect (SA) | Reference architectures and integration patterns across many accounts; pre-sales fit for complex deals | Deep per-customer build | Architecture diagram, RFP answers, reference design |
| Forward deployed engineer (FDE) | The customer’s technical outcome: discovery, deployment hypothesis, POC, migration, first-line diagnosis, translating customer problems into internal tickets; often writes code in the customer’s repo | Engine internals, model research, pricing authority | Discovery notes, POC plan, benchmark report, migration runbook, escalation report |
| MLE (customer-facing or platform) | Fine-tuning runs, eval pipelines, model packaging, quality debugging | Contract, account relationship | Training config, eval report, model card |
| MTS / performance engineer | Engine, kernels, scheduler, quantization recipes, per-model tuning | Customer relationship, POC scope | Engine flags, kernel patch, perf regression fix |
| Research scientist | New methods: speculators, quantization schemes, post-training recipes | Production SLOs | Paper, prototype, recipe |
| Account executive (AE) | Commercial relationship, pricing, contract, renewal | Technical claims | Order form, mutual action plan |
Titles overlap between companies. At some, the FDE, SE, and SA are one person; at others, the FDE writes production code and the SE never leaves pre-sales. Ask in the interview which rows the role covers. The topics below mark which role usually leads each activity.
Where the depth lives
Section titled “Where the depth lives”| Customer says | Where to get the depth |
|---|---|
| “Why is DeepSeek cheaper to serve than a dense model of the same size?” | Foundation Models (MoE, MLA), LLM Systems |
| “Our fine-tune works in Transformers but outputs garbage on your endpoint.” | Model Loading (chat templates, tokenizers), 05 |
| “How many GPUs do we need for 50 req/s?” | Serving & Load, 02 |
| “Latency spikes every few minutes.” | Serving & Load (chunked prefill, preemption), 07 |
| “Should we do LoRA or full fine-tuning? Or RL?” | Training & Post-Training |
| “Is it cheaper than running it ourselves?” | 06 |
| “Why should we trust an open model over GPT or Claude?” | 04 (eval parity on their data) |
Quick Start
Section titled “Quick Start”- New to customer-facing work? Start with 01: the question bank and the note template are the whole job on day one.
- Asked “how many GPUs” on a call? 02 has the worked sizing example and the serverless versus dedicated decision.
- Customer wants a benchmark? 03: run it on their traffic shape, from their region, against their SLO.
- Running a POC? 04 for the plan, then 05 for the migration that follows a successful one.
- Procurement or the CISO just joined the thread? 06.
- Something is broken and it is not yours to fix? 07.
See Study Plan for the schedule.