136 Million PUTs for 17 GiB of Data

clickhouse s3 finops observability mechanism

The S3 line in one of our AWS accounts had been under a euro a month all year. Then it was $733 in nine days. This is the account we run our observability stack in, and the ClickHouse cluster sitting on top of that bucket referenced 17 GiB of data on it.

LineCostQuantity
PUT, COPY, POST, LIST requests$662136 million
Storage$582.6 TB-months
GET requests$1538 million

Storage was eight percent of the bill. This was not a storage problem, and the estimate that said the tier would cost $66 a month had priced the wrong unit.

What a part costs on S3

A ClickHouse MergeTree table stores its data as parts, and a part is a directory. Inside it there is one data file and one mark file per column, plus a handful of metadata files. On a local disk that is a directory listing and nobody thinks about it. On an S3 disk every one of those files is an object, and every time ClickHouse writes a part, each file is a PUT.

Our traces table has 79 columns. One part is about 160 objects. Write a day’s worth of spans as a few hundred whole parts and the request count is a rounding error. Write the same bytes as 90 million fragments and it is the whole bill. The data is identical either way; the number of times you asked S3 to accept it is not.

This is the thing the pricing page does not help with. It quotes cents per GB-month, and a tier holding 2.6 TB at that rate is $66. What it charges you for is operations, and the operation count is a property of how the engine writes, not how much it writes. We had two settings making it write badly, and neither was visible in the values file.

Defect one: the batch that never filled

The OpenTelemetry collectors that feed ClickHouse use a batch processor, and the chart sets it to flush at 50,000 rows or after one second, whichever comes first. At our volume the row ceiling never came first. The timeout did, every second, on every collector, so each collector issued one INSERT per second and each INSERT became a part.

Measured on one shard: 854,554 parts created in a day, at an average of 9 KiB each. Raising the timeout to 30 seconds took the average part to 14.5 MiB. Same data, about 1,600 times fewer objects. A side effect worth naming: the small parts had also been tripping the parts-per-partition insert limit, so ingest had been rejecting batches on and off for the same reason the bill was high. It was an availability defect that happened to also be a cost defect.

Two things made this hard to see. The first is that a chart default is indistinguishable from “unset” in a values diff. Nobody had written timeout: 1s anywhere, so no diff ever showed it, and it survived every chart bump. The second is worse. Our own config notes recorded the one-second setting as a deliberate choice, matched to a lower environment, and dismissed the tighter value on the production instance as legacy cruft. A wrong decision that has been written down is harder to see than one that has not, because the note answers the question before anyone asks it.

Defect two: the flag that also switched off retention

The chart sets prefer_not_to_merge = 1 on the S3 volume of the storage policy. The intent is reasonable. Merging cold data means reading and rewriting it, and on S3 every rewrite is more requests, so the default says leave cold parts alone.

On the ClickHouse version range we run, that flag does a second thing. Deletion by TTL is implemented as a merge, and the merge selector filters out parts on volumes marked with the flag, so DELETE TTL never runs there. Parts moved to S3, were never compacted, and were never expired. Retention was unreachable by construction.

The object count told the story once we looked at it: 6 million, 13 million, 44 million, 90 million, and it never once went down. The query that settled it was a one-liner against system.parts, grouped by disk name: ClickHouse referenced 17 GiB on the S3 disk. The bucket held 48 TB. Three orders of magnitude of objects that the database had already forgotten about and that nothing was going to delete.

One more thing the same query surfaced. The traces table had 185 MiB of stale remnant on the cold tier, while logs and metrics had tiered correctly. So one signal had lost its cold data entirely and the tier as a whole looked healthy. A cold tier can be silently broken for one table and fine for the rest.

690x

There was a discussion in the middle of this about whether S3 is actually cheaper than EBS for cold data, and the number that ended it came from Cost Explorer. On the worst day the bucket grew by 33 TB. Real ingest that day was under 50 GiB. That is roughly 690 times write amplification, and it is not a pricing problem. S3 is cheaper than EBS per byte, by a lot. It is a merge-loop problem, and no per-byte price survives being multiplied by 690.

The general version is this. Merges are what turn many small parts into few large ones. So every cause of starved merges converts directly into request spend the moment tiering is on: a hot disk with no headroom left to merge into, a merge-suppressing flag on the cold volume, a batch rate that outruns consolidation. We normally file those as ingest-health concerns. Under a tiered storage policy they are cost concerns with the same root cause.

The corollary is one I would have argued against a month earlier. Do not trim hot-disk headroom to buy a longer retention window. Reclaiming gp3 at a few cents per GB-month while risking request charges two orders of magnitude higher is a trade you lose every time.

The fix, and the one nobody committed

Two settings closed it. The batch timeout, from one second to thirty. And the merge flag, from one to zero on the S3 volume, so DELETE TTL runs at all. The cost of allowing merges on cold data turned out to be about $3 a month, which is not the argument. The argument is asymmetry: with merges suppressed, one producer drifting past the hot window costs hundreds of thousands of PUTs a day, permanently; with merges allowed, those parts compact within hours.

The third piece is the one now carrying the whole defence: a hot window long enough that parts finish merging on local disk before they move. Fully merged, a day partition on one shard settles to a couple of parts, so the cluster moves about 15 parts a day, around 2,500 PUTs. Unmerged, the same day’s churn across the cluster was 75,000 parts, which at 160 objects each is 12 million PUTs from one table. The difference between those two numbers is entirely whether merging finished before the move.

Then the part I want to be honest about. The 30-second batch timeout that produced the 1,600x reduction had been applied by hand to the live Helm release during the incident and never committed. It sat there with four other hand-applied values. When we ran the infrastructure diff before touching the tier again, the tool reported the release changing from “failed” to “deployed” and nothing else, with a line saying dozens of attributes were unchanged. It compares against its own state, not the cluster. A routine apply would have reverted the fix without a word of warning and re-run the incident on top of whatever else that apply was for.

A fix applied live under pressure is not done until the source reproduces it. The check is mechanical: render the values from source, pull the values the cluster actually holds, and diff them structurally. If those two disagree, the thing that saved you exists only in memory.

Where this generalises, and the trap

Every tiered store built on object storage has this shape. Loki chunks, Thanos and Cortex blocks, Iceberg and Delta small files before compaction, Druid segments in deep storage. Object storage prices operations, so the cost is a function of how many things you write and rewrite. Engines built on immutable parts exist to rewrite things; that is what a merge is. Put the two together and any setting that changes how often the engine writes is a cost setting, whether or not it is labeled as one.

Before enabling a cold tier, measure parts per day on the hot tier. Multiply by files per part. Price the requests. Then, once it is running, check that the tier actually deletes, because a retention setting that cannot execute is indistinguishable from one that is working until the bucket is 60 TB.

The trap is estimating a cold tier by GB-months. That number is always small, it is always the one on the pricing page, and it is the one thing this bill was not about.

$ cat GIT .md
· 7 min read

84 Repositories Vanished. The Fix Was mkdir.

S3 has no empty directories, and in a fully packed git repository the refs directories are exactly that: empty. A file-by-file sync carried every object intact and dropped the two directories git requires to call something a repository, so 84 of 517 came back unreadable while the database still said they had commits. The migration's health checks stayed green the whole time, because none of them ever open a repository.

git aws s3 gitlab mechanism
$ cat KUBERNETES .md
· 7 min read

The Constraint Was Satisfied. The Zone Was Empty Anyway.

A topology spread constraint is only ever a statement about the population its selector matches, and an operator-set label can quietly pool sibling Deployments into one shared count. Three collectors each satisfied their hard zone spread while the union of them left a zone empty, and the standard repair converged on the same wrong answer every time.

kubernetes scheduling opentelemetry finops mechanism
$ cat KAFKA .md
· 8 min read

The Error Was 25 Hours Old. The Client Was Healthy the Whole Time.

A Kafka client wrapper that recovers by counting errors and panicking past a threshold is recovery proportional to traffic: stream processors trip it in seconds, quiet request-driven producers never do, and so they latch the last error and serve it indefinitely while every health signal stays green. The tell is in the error's own digits: an elapsed-time value that is byte-identical across occurrences is one cached event, not a recurring failure.

kafka resilience mechanism reliability observability