September 18, 2026 · 7 min read
84 Repositories Vanished. The Fix Was mkdir.
Three days after we migrated a self-hosted GitLab instance, a colleague asked whether two projects had disappeared. The web UI showed the project pages, the issue lists, the members. The database still recorded 125 commits and 584 KB for one of them. Asking git for anything returned “A repository for this project does not exist yet.”
It was not two projects. It was 84 of 517 across the instance, 34 of them active, including production services. Nobody had noticed for three days, because the affected repositories were, almost by definition, the ones nobody had touched recently.
The signature that says nobody deleted anything
The first instinct on hearing “repositories vanished” is to look for a deletion. The data said otherwise, and the split was clean enough to be a rule.
Every project record was intact: issues, merge requests, members, events, a correct commit count, a correct size in the database. The git layer underneath was absent. When metadata survives and data does not, deletion is the wrong hypothesis; something in the storage layer moved wrong. Metadata gone too would mean a deletion. Metadata alive with data missing means a transfer or a mount, and points you at the path the bytes took rather than at anyone’s actions.
That split kept the investigation from wasting its first hour on who did what, and pointed it at the migration instead.
What git actually requires
A git repository is not “the files”. It is a graph of objects, plus a set of references pointing into that graph. The references live under refs/, with branch tips in refs/heads and tags in refs/tags, traditionally as one small file per reference.
As repositories grow, housekeeping packs those references into a single packed-refs file, the same way loose objects get packed. It is a routine performance optimisation, invisible to every user of the repository. But once packing is complete, refs/heads and refs/tags contain no files. They are empty directories.
Git still requires refs/ to exist before it will call a directory a repository. So the repositories that had been maintained most thoroughly, the packed ones, were exactly the ones where the directory git demands was empty.
S3 has no empty directories
The migration had staged the Gitaly storage tree through S3 and restored it with aws s3 sync. That path had been chosen for good reasons: no routing changes, no security-group changes, nothing mutated on the network path, and a sync that is incremental in both directions so the pre-seed-then-delta shape survives a short freeze window.
But S3 keys are flat. There are no directories in the object model at all; a directory is an inference a client draws from key prefixes, and a directory with no files under it has nothing to infer it from. It does not arrive.
Nothing about this is visible in a byte count. We had compared sizes on both sides, and both sides agreed, because empty directories hold no bytes. Every object transferred intact. The two empty directories that git requires simply did not exist at the destination, and git refused to recognise the tree as a repository, with every object of every commit sitting right there.
This is the part worth internalising. The copy was faithful. It carried exactly what its contract says it carries: objects. The contract is not the filesystem, and the filesystem carried the meaning.
The arithmetic that turned a guess into a root cause
For a while the selection looked random, which is the shape that makes you doubt the hypothesis. Why these 84? Some were years idle, but one had been pushed to six weeks ago. Randomness is what a wrong theory looks like from inside.
Then we counted. 517 repositories with content. 433 of them had loose reference files, so their refs/ directories contained files and therefore survived the sync. 84 did not. 517 equals 433 plus 84, with no remainder.
The damage was perfectly determined by whether housekeeping had packed the repository. That is why it skewed old and idle, and still caught something recent: packing is a function of history, not of current activity. A correlation that survives being counted exactly is a root cause; the remainder is the proof.
The checks that stayed green
Here is the uncomfortable part. The migration’s exit criteria had been signed off, and the checks named in them had passed. Every health signal was green for all three days.
| Check | What it actually does |
|---|---|
gitlab:gitaly:check | One line: probes the service, never opens a repository |
gitlab:check | ~560 lines of configuration, connectivity, namespaces |
The misleading entry is gitlab:check. It prints one line per project, which looks for all the world like a per-repository check. It is filed under “Projects have namespace”. It never touches a repository. Both checks returned clean with 16 percent of the instance’s repositories unreadable, and both would return clean again tomorrow under the same fault.
The tasks that do check are git:fsck, which verifies integrity across every repository, and git:checksum_projects, which checksums every project’s references. Neither is part of gitlab:check, which is exactly why neither ran. The exit criteria named the two tasks that provably could not detect the failure they were meant to gate.
checksum_projects is the right tool around any migration, and it needs no knowledge of what went wrong to catch it: run it on the source before the move, on the target after, and diff. A dropped refs/ shows up immediately. Verification only proves what it actually tests, and the cheapest time to learn what a check tests is before you depend on it.
Where this generalises, and the blind spot
Recovery was four mkdir calls per repository. No restore, no restarts, no data loss at any point. They came back the instant the directories existed.
The rules, in order of how much they cost us:
Empty directories are data. Any tree whose structure carries meaning, a git storage root, a mail spool, an application directory with a lock directory, has to be moved with a tool whose contract includes the filesystem. tar the tree into S3 rather than syncing file by file and the entire problem disappears, because tar carries the directories as entries rather than inferring them from keys.
A green check is only evidence about the question that check asks. A migration exit criteria is a list of questions, and if nobody checked what the checks actually test, the list may be a list of the wrong questions.
And the honest one. Before this migration, the sync path had been studied and an enumeration of what it silently drops was written down: hard links, symlinks, ownership. Each item measured carefully at the time. Empty directories were not on the list. The note was written during the migration that later used this exact path, and the one item missing from an otherwise careful enumeration is the one that broke 84 repositories.
An enumeration of what a tool drops is only as good as its blind spots, and the edges are invisible from inside, before and after in exactly the same way. The check that would have caught it is one command, and it belongs in front of any object-storage move: find <tree> -type d -empty | wc -l. If that number is zero, sync can carry the tree. If it is not, the count is the list of things about to go missing, written down before they go rather than after.