The Registry Answered in 45 Milliseconds. The Build Took Two Minutes Longer.

gitlab debugging migration reliability observability

A few weeks after we relocated our self-hosted GitLab instance to a new account and a new host, the solutions team reported that npm installs against the group package registry had got slow. Builds that took 1:40 to 2:00 now took about 3:30. A clean-cache npm install --verbose showed the shape of it: repeated 500s, each followed by eventual success. The timing lined up with the migration, and the regression was consistent, so the report came in as “slow since the relocation.” It was a reasonable report. It was also wrong in a specific and instructive way.

A constant added interval is a client retry, not a resource problem

The first thing I timed was the failing request itself: 500, in 45 milliseconds. Every single one.

That is the whole diagnosis in one number. A resource problem, a slow disk, a starved CPU, an undersized instance, produces latency that varies under load and degrades gradually. A constant added interval is something else: a fixed pause inserted by the client between attempts. npm’s fetch retry policy walks a backoff schedule, roughly T, T+10 seconds, T+70 seconds, before giving up on an endpoint. Add those pauses to the requests that had to be retried and you get almost exactly the reported extra 110 seconds. The server was never slow. The builds were slow because the client was waiting politely between failures.

I could have spent an hour on gp3 IOPS and instance sizing and found nothing, because the disk and the CPU were fine. The lesson generalizes past npm: when “slower since X” resolves to a constant interval rather than a distribution shift, stop measuring the server and start reading the client’s retry configuration. Status codes first, then timings, then dashboards.

One row in 63,000

Every 500 carried the same exception:

RuntimeError: Object Storage is not enabled for Packages::Npm::MetadataCacheUploader

GitLab caches each package’s npm metadata, the packument, in a table called packages_npm_metadata_caches. Like everything GitLab stores through CarrierWave, each row carries a store column: 1 means local disk, 2 means object storage. This instance had object storage globally disabled. One row in that table said file_store = 2, so any request for that package’s metadata asked CarrierWave to fetch from a backend that did not exist, and the request died with a 500.

We had attempted object storage once, in July, and rolled it back. The rollback was thorough in one specific way and blind in another, which is the next section. What matters here is the audit: I queried information_schema for every table carrying a store column, checked all of them, 62 tables and roughly 63,000 rows, and found exactly one row pointing at the remote store. That single row was the entire incident. The 554 actual package files were all local, which is why tarball downloads returned 200 in 60 milliseconds throughout: the registry genuinely was healthy. Only the metadata cache read path was broken, and only for the packages whose cached row had been flipped.

The check existed; the list did not

The uncomfortable part is that this failure was already documented. The rollback procedure for the object storage attempt, written in July, ends with the correct assertion: verify object storage is disabled and query for zero rows at file_store = 2. The check existed. It was even the right check.

It failed because the step above it iterated a hand-written list of model names. Whoever wrote the runbook enumerated the tables they remembered storing files: uploads, artifacts, package files, and so on. Packages::PackageFile was on the list. Packages::Npm::MetadataCache was not, because in July nobody thought of the metadata cache as a file store, so it was never asserted, so its one flipped row survived the rollback pointing into a bucket that was about to be deleted.

A check that exists but does not cover the thing that breaks is worse than no check, because it buys confidence it has not earned. The fix is mechanical: assertions about coverage must enumerate from the schema, not from memory. One information_schema query finds every candidate table regardless of what anyone recalled while writing the runbook. If your verification step contains a literal list of table names, that list is a guess, and it should be treated as one.

Dating the landmine

Attributing the outage took three passes, and the first two were wrong in opposite directions.

The row predated the September cutover by seven weeks, and no object storage bucket existed in either account, so my first answer was “probably broken before the move.” Then the background job returned the full row, showing last_downloaded_at three minutes after updated_at. That timestamp is written only on the success path, so the cache had demonstrably been served at some point, and I corrected to “the move is the likely trigger.”

The journal settled it. The object storage experiment and its revert are dated July 24. The row’s last successful read was July 22, two days earlier, while it was still local. So the sequence was: the July migration flipped the row to remote, the rollback missed it, the bucket was deleted, and the row sat inert for seven weeks because a file_store = 2 row only detonates when something reads that specific object. Nobody pulled that package until this week. The September relocation was innocent; it merely carried a database that already contained the landmine. The team’s correlation was honest and wrong, and I reinforced it before checking.

One detail deserves its own sentence: the store flip was done with update_column, which skips timestamps, so updated_at on the row points at unrelated activity and will actively mislead anyone dating the change from the row alone. When a timestamp matters, check how the column was written before trusting it.

The dangerous part was deleting it

The fix was one row, so the fix was one delete. That delete had its own trap: CarrierWave models fire a removal callback on destroy, and that callback tries to delete the file from its configured backend, which is the operation that raises the error we were trying to clear. The row had to be removed with delete_all, which skips callbacks entirely.

Three layers of backup went first: an RDS snapshot, which is what actually holds the row, snapshots of both EBS volumes, and a dump of the row itself with a ready INSERT statement. All verified before touching anything. The API reads the cache with metadata_cache&.file, so a missing row short-circuits to the regeneration path: the packument is rebuilt from the package files and stored fresh.

The verification was the satisfying part. The regenerated cache landed at file_store = 1 on local disk, at exactly 6,786 bytes, identical in size to the July cache, which is good evidence the content round-tripped rather than merely stopped erroring. Instance-wide afterwards: zero 5xx, zero rows at the remote store, builds back under two minutes.

What to keep

Two rules survived this one. When a regression shows up as a constant added interval, read the client’s retry behavior before touching the server’s resources; a 500 in 45 milliseconds followed by a two-minute delay is the client walking its backoff, not the server struggling. And when a verification step enumerates coverage from memory, replace the list with a schema query; the table nobody remembered is where the next incident is already sitting.

$ cat KAFKA .md
· 8 min read

The Error Was 25 Hours Old. The Client Was Healthy the Whole Time.

A Kafka client wrapper that recovers by counting errors and panicking past a threshold is recovery proportional to traffic: stream processors trip it in seconds, quiet request-driven producers never do, and so they latch the last error and serve it indefinitely while every health signal stays green. The tell is in the error's own digits: an elapsed-time value that is byte-identical across occurrences is one cached event, not a recurring failure.

kafka resilience mechanism reliability observability
$ cat GIT .md
· 7 min read

84 Repositories Vanished. The Fix Was mkdir.

S3 has no empty directories, and in a fully packed git repository the refs directories are exactly that: empty. A file-by-file sync carried every object intact and dropped the two directories git requires to call something a repository, so 84 of 517 came back unreadable while the database still said they had commits. The migration's health checks stayed green the whole time, because none of them ever open a repository.

git aws s3 gitlab mechanism
$ cat CLICKHOUSE .md
· 8 min read

136 Million PUTs for 17 GiB of Data

Object storage bills per operation, and a ClickHouse part on an S3 disk is not one object but one per column. So the cost of a cold tier is a function of how many parts exist, not how many bytes they hold, and every setting that starves merges becomes a line on the bill. Two chart defaults did exactly that, and the fix that stopped it had never been committed.

clickhouse s3 finops observability mechanism