Skip to content

Field Engineering

How customer-facing engineers at an inference and fine-tuning cloud turn technical depth into a customer outcome: discovery, sizing, benchmarks run for a customer, POCs, migrations, commercials, security reviews, and escalations.

Prerequisites: Comfort reading a latency percentile and a cost table. The technical depth (prefill and decode, KV cache, load testing math) lives in LLM Systems & Inference and Serving & Load and is linked here, not repeated.

  • Every call and every week of work ends in a written artifact someone else can act on. If it did not produce one, it did not happen.
  • Discovery exists to produce a deployment hypothesis with numbers. A question that cannot change a number is not worth the customer’s time.
  • A benchmark run for a customer is a contract in progress: their traffic shape, their region, their SLO, curves instead of single numbers, and no promise the data does not support.
  • A POC exists to make a decision, so its criteria must be able to fail.
  • Escalations that list what was ruled out get acted on; forwarded complaints get meetings.
  • Work the topics in order on one mock customer, Lexa (brief in 01). Each topic’s exercise produces one document; together they form a full engagement file.
  • When a topic links into an ML track, open it and redo the math yourself. A field engineer who cannot defend a number in front of the customer’s MLE loses the room.
  • Read every template aloud as if on a call. If a line would make you hesitate, you do not yet know what it asks for.
graph LR
    D[01 Discovery] --> Q[02 Qualification and Sizing]
    Q --> P[03 Performance Testing Engagements]
    Q --> C[06 Commercials and Security]
    P --> POC[04 POC and Evaluation]
    D --> POC
    POC --> M[05 Migration and Cutover]
    POC --> C
    P --> E[07 Escalation and Handoff]
    M --> E
    E --> MOCK[08 Mock Engagement]
    C --> MOCK
#TopicPrimary ReferenceTime
01DiscoveryThe Mom Test (Fitzpatrick) + SRE book ch. 4 (free)3-4 days
02Qualification and SizingServing & Load (free, in repo)3-4 days
03Performance Testing Engagementsvllm bench serve (free) + AIPerf (free)4-5 days
04POC and EvaluationLLM Evaluation (free, in repo)3 days
05Migration and CutoverModel Loading: tokenizers and chat templates (free, in repo)3 days
06Commercials and SecurityAICPA SOC 2 (free) + GDPR text (free)3-4 days
07Escalation and HandoffPostmortem Culture (free)2 days
08Mock EngagementTopics 01 to 07 applied to your own platform from the course; the role-path capstonecourse pass 11

Company facts below were checked against primary sources on October 6, 2026. Product surfaces change monthly; recheck before a call.

All three sell GPU time wrapped in software. The wrapper differs: a per-token API, a dedicated deployment, a training job, or raw cluster capacity. A field engineer has to know which wrapper fits a workload and what each costs the customer in control and money.

SurfaceFireworks AITogether AIBaseten
Serverless per-token APIPay-per-token serverless models (quickstart)Serverless inference, OpenAI-compatible (quickstart, product)Model APIs, compatible with OpenAI Chat Completions and Anthropic Messages (beta); tool calling, structured output, JSON mode (Model APIs)
Dedicated / on-demand deploymentOn-demand deployments on dedicated GPUs with autoscaling (on-demand quickstart, why on-demand)Dedicated endpoints, single-tenant GPUs (docs)Dedicated deployments of open or custom models (first model)
Fine-tuningManaged: SFT, DPO, LoRA. Training API: LoRA and full-parameter, custom loops, RL (GRPO and others via SDK). RFT listed (intro)SFT, LoRA (default) or full fine-tuning, preference tuning with DPO (overview, DPO). RL not listed in fine-tuning docs at reviewTraining Jobs and Loops (training, loops). Supported method list not confirmed
Batch / asyncBatch inference (docs)Batch API (docs)Async inference with webhooks on any dedicated deployment, queue up to 72 h (docs). No separate batch product confirmed
Embeddings / rerankEmbeddings and reranking (docs)Embeddings and rerank API; rerank models need a dedicated endpoint (rerank)BEI engine for embeddings and reranking (BEI)
Custom model deployCustom model upload on on-demand deployments (announcement); multi-LoRA serving (blog)Custom model upload (docs); Dedicated Container Inference for your own Docker image (docs)Truss (package model code) and Chains (multi-step, multi-model pipelines) (Chains, Truss CLI)
GPU clustersNot a headline product at review; not confirmedH100, H200, B200, GB200 clusters with storage for training or large batch (docs, product)Not offered as a raw cluster product; not confirmed
BYOC / VPC / self-hostedBring Your Own Cluster on customer Kubernetes, Private Preview, enterprise only, inference only (BYOC); SageMaker as a compute option (blog)Together Enterprise Platform in customer VPC or on-prem (announcement, architecture)Baseten Cloud, single-tenant isolated VPC, self-hosted in customer VPC, hybrid (hosting options)
Engine workFireAttention custom kernels: V3 targets AMD MI300, V4 targets NVFP4 on B200 (V3, V4)Together inference engine, Together Kernel Collection, ATLAS runtime-learning speculator (ATLAS, research)Baseten Inference Stack: Engine-Builder-LLM on TensorRT-LLM, BIS-LLM for MoE, speculative decoding (EAGLE, MTP, n-gram), KV-aware routing (guide, BIS-LLM)

Notes on the table:

  • Vendor speedup numbers (for example FireAttention versus vLLM, ATLAS throughput) are the vendor’s own benchmarks on the vendor’s chosen workload. Treat them as hypotheses to reproduce on the customer’s traffic (03), the same rule as in LLM Serving Platforms.
  • Compliance claims change and are contractual. Fireworks publishes SOC 2 Type II and HIPAA status and a no-logging default for open-model prompts (data security). Together documents its posture at privacy and security. For any customer, get the current report from the trust center, not from a blog (06).
  • Together also sells Sandbox and Managed Storage (products); Baseten documents a Frontier Gateway for model labs (overview). These are outside this track.

Common to all three: an OpenAI-compatible request surface, a serverless tier for evaluation, a dedicated tier for SLOs, fine-tuning feeding deployment, and a proprietary engine layer on top of or beside vLLM, SGLang, or TensorRT-LLM. The differentiator they sell is the engine and the people. Field engineers are part of the people.

RoleOwnsDoes not ownTypical artifact
Sales engineer (SE)Pre-sales technical fit for one deal: demos, discovery, first sizing, security questionnaire answers, technical winPost-sale delivery, engine internalsDiscovery notes, demo, questionnaire answers, initial hypothesis
Solutions architect (SA)Reference architectures and integration patterns across many accounts; pre-sales fit for complex dealsDeep per-customer buildArchitecture diagram, RFP answers, reference design
Forward deployed engineer (FDE)The customer’s technical outcome: discovery, deployment hypothesis, POC, migration, first-line diagnosis, translating customer problems into internal tickets; often writes code in the customer’s repoEngine internals, model research, pricing authorityDiscovery notes, POC plan, benchmark report, migration runbook, escalation report
MLE (customer-facing or platform)Fine-tuning runs, eval pipelines, model packaging, quality debuggingContract, account relationshipTraining config, eval report, model card
MTS / performance engineerEngine, kernels, scheduler, quantization recipes, per-model tuningCustomer relationship, POC scopeEngine flags, kernel patch, perf regression fix
Research scientistNew methods: speculators, quantization schemes, post-training recipesProduction SLOsPaper, prototype, recipe
Account executive (AE)Commercial relationship, pricing, contract, renewalTechnical claimsOrder form, mutual action plan

Titles overlap between companies. At some, the FDE, SE, and SA are one person; at others, the FDE writes production code and the SE never leaves pre-sales. Ask in the interview which rows the role covers. The topics below mark which role usually leads each activity.

Customer saysWhere to get the depth
“Why is DeepSeek cheaper to serve than a dense model of the same size?”Foundation Models (MoE, MLA), LLM Systems
“Our fine-tune works in Transformers but outputs garbage on your endpoint.”Model Loading (chat templates, tokenizers), 05
“How many GPUs do we need for 50 req/s?”Serving & Load, 02
“Latency spikes every few minutes.”Serving & Load (chunked prefill, preemption), 07
“Should we do LoRA or full fine-tuning? Or RL?”Training & Post-Training
“Is it cheaper than running it ourselves?”06
“Why should we trust an open model over GPT or Claude?”04 (eval parity on their data)
  1. New to customer-facing work? Start with 01: the question bank and the note template are the whole job on day one.
  2. Asked “how many GPUs” on a call? 02 has the worked sizing example and the serverless versus dedicated decision.
  3. Customer wants a benchmark? 03: run it on their traffic shape, from their region, against their SLO.
  4. Running a POC? 04 for the plan, then 05 for the migration that follows a successful one.
  5. Procurement or the CISO just joined the thread? 06.
  6. Something is broken and it is not yours to fix? 07.

See Study Plan for the schedule.