Observability & Monitoring

Your customers should not be your monitoring

We cut the distance between something going wrong and somebody knowing. Alerting that fires on what a user feels, traces that say why, and a stack your own team owns afterwards. Start with a free health check of what you run today.

Get the free health check

What we do

Detection

Find it before support does

Alerts written against what a user would notice rather than against a threshold someone picked in 2023. Tied to an SLO, routed to whoever can act, and linked to a runbook. The ones that fire without anyone acting get deleted.

Diagnosis

Answer the question during the incident

A metric that leads to the request that caused it, and a trace that names the slow dependency. The difference between knowing something is wrong and knowing what to roll back, which is where incident time actually goes.

Cost

A bill that tracks your business

Retention, sampling and storage sized against the volume you really produce. Most estates are paying to keep telemetry nobody has queried in a year, and the fix is usually a policy rather than a migration.

Ownership

Your team runs it, not us

Everything lands as code in your repository, with a written handover and a walkthrough for the engineers who own it afterwards. Nothing is licensed from us and nothing stops working if we stop talking.

What actually goes wrong

Almost every team we meet already runs Prometheus. Installing it was never the hard part. What is missing is the distance between something breaking and somebody knowing: an alert written against what a user feels rather than a threshold nobody has revisited since install, a trace that names the slow dependency, retention that still holds last Tuesday when you go looking. Close that distance and an outage stays an incident. Leave it open and you find out from a customer.

What it looked like at Comper

Comper had no observability on their production platform at all, so infrastructure problems went unalerted until somebody happened to notice. We built the full stack on Kubernetes with centralised storage and ingestion from several clusters. Within weeks their own team found that services of theirs were being OOMKilled regularly, container restarts nobody had ever seen.

“Victor helped us out quickly by setting up observability in our Kubernetes clusters. The setup is highly available and Victor was very responsive during the whole process, and also available for quick finetuning and setting up extra alerts after the installation. Highly recommended.”

Jouke, Comper

The stack

What gets installed

Open source throughout, managed as code, running in your own infrastructure. Nothing here is ours to take away.

  • Prometheus

    Metrics collection and the query language your alert rules are written in.

  • Loki

    Log aggregation indexed by label rather than by content, so storage stays proportionate.

  • Tempo

    Distributed tracing, linked from the metric or the log line that raised the question.

  • Pyroscope

    Continuous profiling, for when the service is slow and nothing upstream explains it.

  • Grafana Alloy

    One OpenTelemetry collector on every node, in place of an agent per signal.

  • Thanos or Mimir

    A global view across clusters, with downsampling and long-term storage in S3.

Grafana sits in front of all of it. Which pieces you actually need depends on where the gaps are, and a stack with three components you use beats six you do not.

Engagements

Three ways in

Free

Observability Health Check

A read-only look at what you run today, written up as findings ranked by what would help most.

  • Signal-to-noise and alert hygiene
  • Retention, storage and cost
  • Yours to keep, no obligation
Start here

1 week

Monitoring Baseline

Prometheus, Grafana and alerting installed and tuned on one cluster, owned by your team afterwards.

  • Golden-signal dashboards
  • Paging on real conditions
  • Handover session

3-4 weeks

Production Observability Stack

The full build: metrics, logs, traces and profiles, with long-term storage and alerting tied to SLOs.

  • Grafana Alloy as unified collector
  • Long-term storage with downsampling
  • Documentation and walkthrough

Prices are published on the pricing section of the homepage. Multi-cluster and multi-tenant work is quoted after a scoping call, and advisory runs at EUR 125 per hour.

Frequently asked

We already have Prometheus. What is left to do?
Usually three things: retention, which is where most self-built stacks quietly stop at fifteen days; alerting, which tends to be a set of thresholds nobody has revisited since installation; and the logs and traces that were never wired in, so an incident still ends in kubectl logs. The install is the easy part and it is rarely where the value is.
What does the free health check actually involve?
A read-only look at what you run today, then a written list of findings ordered by what would help most: signal-to-noise, alert hygiene, retention, storage and cost. It takes a short call and access to your dashboards. You keep the findings whether or not we work together, and there is no obligation attached to it.
Does this only work on Kubernetes?
No. Kubernetes is where most of the work happens, but the same collectors run on plain virtual machines and bare metal, and managed cloud services are scraped through their own exporters. A hybrid estate is normal rather than a complication.
Self-hosted or SaaS?
Depends on your volume, your team, and what you are required to keep in-house. Self-hosted is usually cheaper past a certain telemetry volume and costs you the operational burden of running it. We have built and run both, so what you get is an estimate of what each costs in practice rather than a preference. If Datadog is the answer, that has its own page.
How long does it take?
One week for a single cluster monitored properly, three to four weeks for a full production stack with logs, traces, profiles and long-term storage. Multi-cluster work is quoted after a short scoping call, because the answer depends on how much of it already exists.
Do you work with teams outside the Netherlands?
Yes. We work remotely with teams across Europe.

Ready to talk?

Independent observability consultants. We cut the time between something breaking and someone knowing, using Prometheus, Grafana, the LGTM stack and OpenTelemetry. Start with a free health check.

Get in touch