Blog

$ cat GITLAB .md
· 7 min read

The Registry Answered in 45 Milliseconds. The Build Took Two Minutes Longer.

After a server migration, builds against a self-hosted package registry got two minutes slower and every report pointed at the move. Every failing request was a 500 in 45 milliseconds; the entire regression was npm's retry backoff. The root cause was one row out of 63,000 still pointing at an object store that no longer existed, and the check written to catch exactly that iterated a hand-written list of tables and missed it.

gitlab debugging migration reliability observability
$ cat GIT .md
· 7 min read

84 Repositories Vanished. The Fix Was mkdir.

S3 has no empty directories, and in a fully packed git repository the refs directories are exactly that: empty. A file-by-file sync carried every object intact and dropped the two directories git requires to call something a repository, so 84 of 517 came back unreadable while the database still said they had commits. The migration's health checks stayed green the whole time, because none of them ever open a repository.

git aws s3 gitlab mechanism
$ cat KUBERNETES .md
· 7 min read

The Constraint Was Satisfied. The Zone Was Empty Anyway.

A topology spread constraint is only ever a statement about the population its selector matches, and an operator-set label can quietly pool sibling Deployments into one shared count. Three collectors each satisfied their hard zone spread while the union of them left a zone empty, and the standard repair converged on the same wrong answer every time.

kubernetes scheduling opentelemetry finops mechanism
$ cat CLICKHOUSE .md
· 8 min read

136 Million PUTs for 17 GiB of Data

Object storage bills per operation, and a ClickHouse part on an S3 disk is not one object but one per column. So the cost of a cold tier is a function of how many parts exist, not how many bytes they hold, and every setting that starves merges becomes a line on the bill. Two chart defaults did exactly that, and the fix that stopped it had never been committed.

clickhouse s3 finops observability mechanism
$ cat KAFKA .md
· 8 min read

The Error Was 25 Hours Old. The Client Was Healthy the Whole Time.

A Kafka client wrapper that recovers by counting errors and panicking past a threshold is recovery proportional to traffic: stream processors trip it in seconds, quiet request-driven producers never do, and so they latch the last error and serve it indefinitely while every health signal stays green. The tell is in the error's own digits: an elapsed-time value that is byte-identical across occurrences is one cached event, not a recurring failure.

kafka resilience mechanism reliability observability
$ cat DATABASES .md
· 7 min read

LakeBase Didn't Reinvent Postgres. It Moved the fsync.

Databricks LakeBase Postgres is real, wire-compatible Postgres: standard drivers, pgvector, PostGIS, unmodified query engine. The marketing headline is database branching, Git for your database. That is the side effect. The mechanism is that the commit path moved off the machine: a transaction acks when a quorum of safekeepers accepts the WAL record, not when a local SSD flushes. Branching, scale-to-zero, and point-in-time recovery all fall out of that one sentence, and so do the costs.

databases postgres architecture mechanism cloud
$ cat LINUX .md
· 7 min read

Opinionated Linux Stops Being a Contradiction When the Opinions Are Coherent.

Omarchy launched this summer and 37signals is moving the whole company to it on a three-year horizon. The interesting part is not the Linux part; it is that the distribution is the first one I have seen where the personality is the lead feature. The portable rule: the same opinionated default is a feature when its reasoning is visible and a bug when its reasoning is hidden.

linux design opinion product devops
$ cat AI .md
· 6 min read

Speculative Decoding Shipped Because It Doesn't Change the Output.

Liquid AI shipped LFM2.5-DSpark on August 20: three draft models under the lfm1.0 license, with a 2.67× mean speedup on H100 and a 57% function-calling latency reduction on Apple Silicon. The headline number is the least interesting thing about the release. The interesting thing is why this kind of release can ship at all. Speculative decoding is exact under greedy decoding, which means it does not change the output. That property is the whole game.

ai inference speculative-decoding mechanism performance
$ cat OBSERVABILITY .md
· 7 min read

The Old Pipeline Lost ~47k Spans. It Didn't. We Counted the Wrong Thing.

During a parallel-run validation, a side-by-side per-service span count showed the new pipeline losing ~0.2% of spans per service across a 60-minute window. Read as regression in the new pipeline, it would have triggered a rollback. Read correctly, it was the old pipeline ingesting tens of thousands of duplicate rows in a few-second window. The storage layer has no unique constraint on the thing being counted; the unit it exposes as 'count' is not the unit the operator thinks it is.

observability clickhouse verification troubleshooting devops
$ cat DOCKER .md
· 5 min read

The Note Said the Image Was Wrong. The Image Was Right. Three Checks Agreed.

A note in the project's knowledge bundle said the image was amd64-only and might not run on Graviton. Three independent checks agreed. The image was not amd64-only, and a native arm64 build succeeded with zero source changes. A claim that has been verified three times is the most dangerous kind of wrong claim.

docker arm64 verification knowledge-management devops
$ cat KAFKA .md
· 8 min read

I Inferred Three Things About a Live System. Two of Them Weren't True.

Three times in one week I made a claim about a live migration by reading a config file, a name prefix, or a template, and twice the claim was wrong. The artefact and the live system are two views of the same thing, and they can agree for reasons the artefact can't tell you. Only one of the views is the truth.

kafka opentelemetry terraform troubleshooting devops
$ cat TERRAFORM .md
· 9 min read

The Plan Was Green Because Two of Three Views Agreed. The Third One Was the Truth.

A terraform plan came back clean for a managed observability stack after an incident fix had been applied by hand. The fix was live, the fix was in the file, and state alone was behind. A plan is a diff between two views, but the system has three. Reading all three before you plan is what turns the green diff from a guess into a decision.

terraform helm eks state troubleshooting devops
$ cat OPENTELEMETRY .md
· 6 min read

The Config Was Byte-Identical. That Proved the Wrong Thing.

We re-derived a collector config, confirmed it was byte-for-byte identical to the one running in production, and shipped it with a stack of green checks behind it. It crash-looped on startup. Byte-identical is a real property; it just answers a question nobody was asking.

opentelemetry kafka terraform devops troubleshooting
$ cat AWS .md
· 6 min read

Two DNS Layers That Look Identical When They're Both Green

A VPC in one account couldn't resolve names in a private zone owned by another. Peering was healthy, routes were right, the DNS-resolution flag was checked. Every signal was green, and the answer was still NXDOMAIN, because two different DNS layers were wearing the same green light.

aws route53 dns vpc networking troubleshooting
$ cat CDKTF .md
· 6 min read

The Bug I Blamed on the Provider

For weeks a terraform plan showed drift I couldn't fix on an S3 bucket that was already configured correctly. I filed it under 'AWS provider quirk' and worked around it. The provider was innocent. My own code was lying to me before Terraform ever saw it.

cdktf terraform iac typescript devops troubleshooting
$ cat KUBERNETES .md
· 8 min read

The EKS 401 That Wasn't About Credentials

Our CI pipeline was minting a valid EKS token, presenting it cleanly, and getting rejected. The problem was not AWS creds. It was EKS access control, and the failure was hiding in plain sight in the cluster's RBAC.

kubernetes eks aws devops platform-engineering troubleshooting
$ cat KUBERNETES .md
· 6 min read

The Agent Control Plane: Omnigent on Kubernetes

Databricks open-sourced Omnigent, a meta-harness that wraps any agent harness behind a unified API with policies, sessions, and shared state. If you have deployed stateful services to Kubernetes before, the path to self-hosting it is shorter than you think.

kubernetes devops platform-engineering ai-agents
$ cat KUBERNETES .md
· 10 min read

StackGres and Strimzi: Day-Two Lessons From a Real Platform Rebuild

Every operator promises a simpler setup. The real payoff shows up later, when a replica loses its control file or a certificate is three weeks from expiring. Here is what StackGres and Strimzi actually changed about running PostgreSQL and Kafka on Kubernetes.

kubernetes operators postgresql kafka devops platform-engineering
$ cat KUBERNETES .md
· 6 min read

Helm deploy succeeded. Config never applied.

A green CI pipeline while the live gateway serves 404s is not a test failure. It is a silent reload gap. Here is the cross-release checksum pattern that closes it in helmfile.

kubernetes helm helmfile devops platform-engineering