Why time to first byte still matters
Every other metric waits on the first byte. Here is where the milliseconds go and how to get them back.
The metric everyone skips
Time to first byte is the least glamorous number in a performance report. It does not describe anything a visitor sees, it rarely shows up in a design review, and most dashboards bury it under the paint metrics. Yet every one of those paint metrics starts its clock after the first byte arrives. A page cannot render HTML it has not received.
When we audited a dozen production sites last quarter, the median first byte sat at 420 milliseconds on a fast connection. On a mid-range phone over a congested cellular link, that number tripled. No amount of image compression or font subsetting recovers time that was lost before the browser had anything to parse.
Where the time goes
A first byte is the sum of several smaller waits:
- DNS resolution, which is usually cached but costs a full round trip when it is not.
- Connection setup, the TCP handshake plus TLS negotiation, two or three round trips on older stacks.
- The request itself traveling to wherever the response is produced.
- Server think time: database queries, template rendering, calls to other services.
- The response traveling back.
The network legs scale with distance. The think time scales with how much work the server does per request. Both are fixable, but they are fixed differently.
Shortening the distance
The cheapest round trip is the one that ends close to the visitor. Edge networks exist for exactly this reason: they terminate the connection a few milliseconds away, reuse warm connections to the origin, and can answer outright when a response is cacheable. Even when a page must be rendered fresh, terminating TLS at the edge removes a surprising amount of latency because the expensive handshake happens over a short hop.
Shrinking the think time
Server think time is where most teams find their biggest wins. Common culprits include:
- Sequential data fetches that could run in parallel.
- A slow query nobody noticed because it is fast on a laptop with ten rows.
- Rendering a whole page when only a fragment changed.
- Synchronous calls to third-party services in the critical path.
Profiling a single request end to end, with timestamps at each boundary, usually reveals the worst offender within an hour.
Measuring it honestly
Lab tools report a single first byte from a single location. Real visitors are spread across continents and devices. Collect the metric from the field, segment it by region, and look at the 75th percentile rather than the median. The median tells you how the typical visit feels. The tail tells you who is quietly leaving.
A practical target
Aim for a first byte under 200 milliseconds at the 75th percentile for visitors near your edge, and under 600 milliseconds everywhere else. Those numbers leave room in the budget for everything that follows, which is the whole point: the first byte is not the experience, but it is the floor the experience stands on.