Skip to content

Commercials and Security

The numbers and paperwork that decide a deal after the technology works: cost per 1M tokens, total cost of ownership (TCO) of self-hosting versus serverless versus dedicated, the breakeven point, committed-use and reserved capacity, security questionnaires, and how procurement and legal fit into the timeline.

Parent track: Field Engineering. Builds on the sizing in 02 and the measured operating point from 03. The cost formulas are derived in Serving & Load sections 7 and 8.

All prices in this topic are placeholders chosen to show the method. They are not any vendor’s list price. Plug in the current price sheet and the customer’s real rates before showing a number to anyone.

  • Quote cost per 1M tokens at the measured operating point, never at the roofline bound or at saturation.
  • TCO of self-hosting is dominated by utilization and engineering time, not the GPU hourly rate. A cheap GPU at 40% utilization with two engineers on call is not cheap.
  • Dedicated beats serverless above a breakeven volume: GPU $/hr / serverless $ per token. Below it, serverless wins; above it, dedicated wins only if utilization actually stays high.
  • Commit to the floor of the forecast, not the peak. Unused commitment is money the customer may lose.
  • Security and legal review is usually the longest pole. Start it in the first week of the POC, and answer questionnaires from the vendor’s maintained answers, never from memory.
  • Rebuild the TCO table below in a spreadsheet with every input as a cell. Then change utilization from 50% to 80% and engineering from 1.5 FTE to 0.5 FTE and see which moves the answer more.
  • Read GDPR Article 28 and Chapter V headings once. You will not give legal advice, but you must know which document a question is about.
  • Download a public CAIQ from any cloud vendor’s trust center and find the ten questions an inference vendor would find hardest (retention, sub-processors, model training on customer data).

Customers do not buy tokens per second; they buy a cost line and a risk profile. The commercial conversation converts the benchmark into money (cost per 1M tokens, TCO, breakeven, commitment size), and the security conversation converts the deployment architecture into assurances (reports, contracts, data flows). Both are evidence-based, both have owners other than the field engineer, and both take calendar time the technical plan must account for.

Serverless prices are usually per token, often with separate input and output rates. Compute the blended rate from the customer’s own mix:

cost per request = T_in x p_in / 1e6 + T_out x p_out / 1e6
blended $ per 1M = cost per request / (T_in + T_out) x 1e6
support copilot (02): T_in = 2,000, T_out = 300
placeholders: p_in = $0.60, p_out = $2.40 per 1M
cost per request = 2,000 x 0.60e-6 + 300 x 2.40e-6 = $0.00120 + $0.00072 = $0.00192
blended = 0.00192 / 2,300 x 1e6 = ~$0.83 per 1M tokens

Input-heavy workloads (RAG, long documents, contract review) are dominated by the input rate. Some providers discount cached input tokens; whether and how much is vendor-specific (check the current price sheet).

Dedicated cost per 1M tokens at an operating point:

$ per 1M = (GPU $/hr x GPUs) / (tokens/s served x 3600) x 1e6
over a month: $ per 1M = (GPU-hours x $/GPU-hour) / tokens served x 1e6
support copilot: (4,320 GPU-h x $4.00) / 89,000M tokens = ~$0.19 per 1M

The monthly form already includes idle time, which is why it is the honest one.

2. TCO: Self-Hosting versus Serverless versus Dedicated

Section titled “2. TCO: Self-Hosting versus Serverless versus Dedicated”

Self-hosting looks cheapest on a GPU price list because the list omits utilization and people. A TCO includes:

Cost lineSelf-host (own cloud GPUs)Dedicated (vendor)Serverless (vendor)
GPU timeAll reserved hours, used or idleGPU-hours including the autoscaling floorNone (in the token price)
Utilization riskCustomer’sCustomer’s, reduced by autoscalingVendor’s
EngineeringEngine tuning, upgrades, kernels, autoscaling, on-callIntegration and monitoringIntegration
Capacity riskCustomer must secure GPUsVendor capacity, may need reservationRate limits by tier
HiddenObservability, storage, egress, security review of the stackEgress, private networkingRate-limit headroom

Worked example, support copilot (89B tokens/month, 02). Every price and salary is a placeholder.

LineSelf-hostDedicatedServerless
GPU or token cost12 GPUs x 720 h x $2.50 = $21,600 (no cheap autoscaling on reserved GPUs)4,320 GPU-h x $4.00 = $17,28089,000 x $0.90 = $80,100
Utilization of paid GPU-hours~50% (needs ~4,320 of 8,640 h)High while the floor holdsn/a
Engineering1.5 FTE x $250k / 12 = $31,2500.25 FTE = $5,2080.1 FTE = $2,083
Other (observability, storage, egress)$2,000$500$0
Monthly total~$54,850~$22,990~$82,180
Effective $ per 1M tokens~$0.62~$0.26~$0.92

Read the result, not the digits: with these inputs, self-hosting loses to dedicated because of engineering time and idle reserved GPUs, and serverless loses because the volume is well above breakeven. The self-host column also assumes the customer’s engine matches the vendor’s throughput per GPU, which is the vendor’s core claim and should be tested, not assumed (03).

Per GPU: dedicated is cheaper when the tokens a GPU serves per hour exceed what the same money buys on serverless.

breakeven tokens per GPU-hour = R / (P / 1e6)
placeholders R = $4.00 per GPU-hour, P = $0.90 per 1M blended:
4.00 / 0.90e-6 = ~4.44M tokens per GPU-hour = ~1,235 tokens/s sustained per GPU

Per deployment: a dedicated deployment has a minimum monthly cost (its floor). Volume below that floor’s token equivalent belongs on serverless.

floor: one TP4 replica, always on = 4 GPUs x 720 h x $4.00 = $11,520/month
breakeven monthly tokens = 11,520 / 0.90 x 1e6 = ~12.8B tokens/month (~4,900 tok/s average)
support copilot: 89B tokens/month -> well above breakeven -> dedicated

Serving & Load section 7 works the same formula with different placeholders. Two checks before trusting a breakeven:

  • The dedicated deployment must actually serve that volume within the SLO (goodput from 03), not just at saturation.
  • The traffic curve matters: a workload with a high peak-to-mean ratio pays for idle GPUs unless autoscaling is fast enough for the ramp.
ConceptWhat the customer getsWhat the customer risksField rule
On-demandPay per hour or per token, no commitmentPrice and availability not guaranteedUse for evaluation and POCs
Reserved capacitySpecific GPUs held for a term at a lower ratePays whether used or notSize to the measured floor, not the peak
Committed spendA discount in exchange for spending an agreed amount over a term, often across productsUnused commitment may be forfeited (terms vary by vendor and contract; not confirmed for any specific vendor)Commit to the floor of the forecast; leave growth to on-demand
Ramp dealCommitment that grows by quarter as migration proceedsMigration slips while the commitment risesTie ramp steps to migration milestones in 05

The AE owns price and terms. The field engineer owns the inputs: measured operating point, forecast floor and peak, migration schedule, and the risk list. Never quote a discount.

ItemWhat it isWhat to provideSource
SOC 2 Type IIIndependent CPA attestation of a service organization’s controls against the AICPA Trust Services Criteria: security, availability, processing integrity, confidentiality, privacy. Type II covers whether controls operated effectively over a period (commonly several months to a year); Type I covers design at a point in timeThe current report from the vendor’s trust center, usually under NDA; a bridge letter if the report period ended months agoAICPA
ISO/IEC 27001Certification of an information security management systemCertificate and scope statementISO
HIPAA BAAA covered entity may disclose protected health information (PHI) to a business associate only with satisfactory written assurances, a contract meeting 45 CFR 164.504(e). HHS guidance treats a cloud provider handling ePHI as a business associate, directly liable for unauthorized uses and for Security Rule safeguardsWhether the vendor signs a BAA, for which products (often not all), and on which tiers45 CFR 164.504, HHS cloud guidance
GDPR processor terms (DPA)The vendor processes personal data on the customer’s behalf. Article 28 requires a contract setting the subject matter, duration, nature and purpose of processing, data types, and data subjects; sub-processors need the controller’s prior written authorizationThe vendor’s DPA and sub-processor listArt. 28
International transfersTransfers outside the EU/EEA need a Chapter V basis: an adequacy decision (for example, the EU-US Data Privacy Framework, adopted 10 July 2023, for certified US companies) or safeguards such as standard contractual clausesTransfer mechanism and certification statusChapter V, Commission, DPF
Data residencyProcessing (and any storage) stays in a named regionRegion list for the tier and model; where logs, metrics, and support access liveVendor docs
Zero data retention (ZDR)Prompts and completions are not stored after the requestExactly what is still kept (token counts for billing, abuse-monitoring logs, error traces), for how long, and on which tiersVendor docs and contract
No training on customer dataCustomer prompts and outputs are not used to train vendor modelsContract clause, not a blog postContract
VPC / BYOCPrivate networking to a vendor deployment, or the deployment runs in the customer’s cloud accountNetwork diagram, who operates it, who holds keys, which features are missing in that modeVendor docs (role landscape)
QuestionnairePublisherShape
CAIQ v4Cloud Security AllianceAbout 260 yes/no questions aligned with the Cloud Controls Matrix, including the shared security responsibility model
SIGShared AssessmentsBroad third-party risk questionnaire for any vendor, common in financial services and healthcare
Custom spreadsheetThe customer’s security teamOften a mix of the above plus AI-specific questions: training on data, retention, model provenance, prompt injection handling

Rules: answer from the vendor’s maintained answer library and current reports; route any new question to the security team with a due date; never improvise an answer about retention, sub-processors, or encryption. A wrong “yes” in a questionnaire can become a contractual misrepresentation.

flowchart LR
    subgraph Technical
        D["Discovery"] --> H["Hypothesis"] --> POC["POC"] --> R["Readout"]
    end
    subgraph Commercial["Security, legal, procurement"]
        NDA["NDA"] --> SQ["Security questionnaire<br/>and SOC 2 report"]
        SQ --> DPA["DPA / BAA<br/>redlines"]
        DPA --> MSA["MSA and order form"]
        MSA --> PO["Purchase order"]
    end
    D -.->|"start in week 1"| NDA
    SQ -.->|"needed before production data"| POC
    R --> MSA
    PO --> PROD["Production access"]

Key ideas:

  • Run both tracks in parallel. Discovery should surface who owns security review and how long it takes (01); start it the day production data is mentioned.
  • Production data needs paperwork first. A POC on real prompts usually needs the NDA, the security review, and the DPA (or BAA) signed.
  • Redlines take calendar time. DPA and BAA terms go between legal teams; the field engineer supplies accurate technical descriptions of data flows and does not negotiate terms.
  • Procurement has its own gates: vendor onboarding, insurance certificates, payment terms. Ask the customer’s champion for their checklist early.
  • How long each step takes varies widely by customer size and industry; this track gives no typical durations (not sourced). Ask the customer.
ConceptConnected TrackApplication
Cost per token, breakeven, autoscaling floorServing & LoadSections 1 to 3
Authorization, tenancy, access controlAuthorization & Access ControlQuestionnaire answers on access
Private networking, VPCs, KubernetesCloud NativeVPC and BYOC architecture
Audit logging, retentionObservabilityWhat ZDR means for logs and traces
CompanyHow This AppearsDifficulty
Fireworks AI / Together AI / BasetenTCO and breakeven conversations; trust center and questionnaire supportIntermediate
AWS / GCP / AzureCommitted use and reserved capacity programs; BAA-eligible service listsIntermediate
OpenAI / Anthropic (enterprise)ZDR, data residency, and DPA conversations for API customersIntermediate
Healthcare and legal-tech customersBAA and residency as gating requirementsAdvanced

Deliverable: a commercial and security brief for Lexa (brief), two pages, built on your 02 hypothesis and 03 operating point:

  1. Blended cost per 1M tokens for their mix on serverless (placeholder prices, labelled), and dedicated cost per 1M at your operating point.
  2. A TCO table for self-host, dedicated, and serverless, with the online path and the nightly job as separate rows, compared with their $60k/month spend and the “cut by half” goal.
  3. The breakeven volume and where Lexa sits today and after six months of 15% monthly growth.
  4. A commitment recommendation (floor, ramp) with the risk to Lexa if migration slips.
  5. The CISO meeting: the list of documents you bring (report, DPA, sub-processor list, residency and retention statement) and the five questions you expect, each with who answers it.
  • Every price is labelled as a placeholder or sourced with a date.
  • The nightly job is priced on a batch or async surface.
  • The growth projection changes or confirms the tier choice, and the brief says which.
  • No security answer is asserted without naming the document that supports it.