Field Notes
All posts

Observability for small teams

You do not need a platform team to know when your site is broken. Three signals cover most of it.

Start with the questions

Observability tools can collect almost anything, which makes it easy to collect everything and understand nothing. A small team does better starting from the questions it needs answered:

  • Is the site up right now?
  • Is it slower than usual?
  • Are visitors hitting errors?
  • When something broke, what changed?

Everything you instrument should help answer one of these.

Signal one: synthetic checks

A synthetic check requests a page from outside your infrastructure every minute and records whether it succeeded and how long it took. It is the simplest possible answer to "is it up", and it catches failures that internal health checks miss, such as an expired certificate or a misconfigured DNS record.

Check the pages that matter most, not just the home page. A home page served from cache can look healthy while every page that hits the database is failing.

Signal two: real user metrics

Synthetic checks run from a handful of locations on fast connections. Real user monitoring collects timing and errors from actual visitors. It shows how the site performs on real devices across real networks, and it reveals regional problems a synthetic check would never see.

Signal three: structured logs

Logs are most useful when every line carries the same core fields: a timestamp, a request ID, the route, the status, and the duration. With those, you can answer most questions with a simple query. Free-form log messages are much harder to search under pressure.

Alert on symptoms

Alerts should fire when visitors are affected, not when a single internal metric twitches. An elevated error rate on real requests is a good alert. High CPU on one machine usually is not, unless it is causing errors. Every alert that fires without requiring action trains people to ignore alerts.

Record changes

Most incidents follow a change: a deploy, a configuration edit, a dependency update. Mark deploys on your dashboards so that a spike in errors can be lined up with the release that preceded it. This single habit shortens investigations more than any tool.

Watch the watchers

A monitor that silently stops reporting looks exactly like a healthy system. Make sure your checks alert when they themselves fail to run, and periodically break something deliberately in a safe environment to confirm the alert fires.

Grow as needed

Tracing, profiling and custom metrics are valuable, but add them when a real question demands them. Three well understood signals beat thirty half-configured ones.