Schemaless, Not Structureless: A True Columnar Index for Log Data
Schemaless, Not Structureless: A True Columnar Index for Log Data
The problem: logs have no fixed shape
Logs are the wild west of telemetry. Every service logs differently, every team adds its own fields, and enrichment pipelines bolt on more — pod names, regions, build versions — long after the code was written. Each log line arrives as a loose bag of facets — key-value attributes — and no two lines carry quite the same bag. For decades, storage engines have forced an uncomfortable choice on this data: keep it flexible and query it slowly, or make it fast and lose the flexibility.
We built a storage index that removes that tradeoff: a schemaless index that gives every facet — whether it appears on a billion log lines or just three — a true, typed, compressed column of its own, created automatically at ingest time. This post explains how it works, why it is fundamentally faster than JSON-blob and row-oriented designs, and why it can keep absorbing new fields without ever asking for a schema migration. The design applies to any telemetry — traces, events, user sessions — but logs are where schema chaos is most extreme, so logs are the lens we'll use.
The problem: logs have no fixed shape
Consider four log lines from a single application fleet:
Three things make this data hard to store well:
- The key space is unbounded. Across a fleet of services, teams, and deploys, the set of distinct facets runs into the thousands — and nobody can enumerate it up front.
- Each line is sparse. Any single log line carries only a handful of those thousands of keys. Most keys are absent from most lines.
- The shape changes constantly. Tomorrow's deploy starts logging
checkout.retries; next week's addsfeature.flag.variant. New keys appear daily, without warning or coordination.
The usual answers, and what they cost
Three architectural patterns dominate how analytical systems store log facets today.
1. The JSON blob. Store each line's facets as one serialized JSON string in a single column. Ingest is trivial and nothing is ever dropped — but every query pays the full price. Asking "what is the average duration_ms for the payments service?" means reading every blob, parsing every blob, and extracting one field from each, even though the other 99% of each blob is irrelevant to the question. Compression suffers too.
2. Row-oriented document storage. Store each log line as a document, retrieved whole. This is excellent for "show me this one log line" — and structurally wrong for analytics. Faceting on a single key across ten million lines means reading ten million whole documents. Such systems stay fast only by building inverted indexes on everything in advance, multiplying ingest cost and storage for facets that may never be queried.
3. The rigid columnar schema. Classic columnar storage delivers exactly the read pattern log analytics wants: scan only the columns the query touches. But it demands the schema up front. Every new log field is a schema migration; every undeclared field is either dropped or stuffed into… a JSON blob, returning us to pattern 1.
| JSON blob | Row/document store | Rigid columnar | Schemaless index | |
|---|---|---|---|---|
| Accepts any shape | ✅ | ✅ | ❌ | ✅ |
| Reads only what the query needs | ❌ | ❌ | ✅ | ✅ |
| Columnar compression | ❌ | ❌ | ✅ | ✅ |
| New key requires migration | No | No | Yes | No |
| Values stored typed (no re-parsing) | ❌ | partly | ✅ | ✅ |
The fourth column is the subject of the rest of this post.
The schemaless index: every key becomes a real column
The core idea is simple to state: decompose log lines at ingest time, so that every distinct facet gets its own typed, sparse column — automatically. Flexibility is preserved at the boundary where data arrives; full columnar structure exists everywhere data is stored and read.
Four structures make this work:
A key dictionary. Every distinct key, together with its inferred value type, gets one entry in a compact dictionary — the index's table of contents.
A sparse column per key. Each key's values are stored contiguously, in arrival order — but only for the lines that actually have the key.
A presence map per key. A compact bitmap records exactly which lines carry the key. This is the bridge between "log line #7" and "the 3rd value in this column".
Typed values. Types are inferred per key at ingest: numbers are stored as numbers, strings as strings, booleans as booleans.
Why this is fast
Each speed advantage falls directly out of a structure above — no benchmarks required to see why.
I/O proportional to the question, not the data. A query touching one key out of 200 reads roughly 1/200th of the facet data. The blob design reads and parses 100% of it to answer the same question.
No parsing, ever. Values are stored already decoded and typed. Reading an integer is reading an integer — not tokenizing a JSON string.
Compression that actually works. A column of 10,000 status codes is highly repetitive — exactly what block compressors are built for.
Encodings tailored to each key. Per-key columns unlock per-key encoding choices.
Sparsity is free. Rare keys cost almost nothing — a key on 1% of lines uses ~1% of the storage of a dense column, plus a tiny presence map.
Built to grow
Extensibility is not a feature added on top of this design — it is the design.
The data model evolves itself. New keys become columns automatically; new value types coexist as sibling columns.
Storage decisions are made per key. Compression codec and presence-map strategy are chosen per column.
The key is the unit of optimization. Because every facet already lives in its own column, future accelerations attach naturally at the same granularity.
Conclusion
The flexibility-versus-performance tradeoff for log data was never a law of nature. It was an artifact of one decision: treating a flexible log line as an opaque blob. Move the decomposition to ingest — a dictionary of keys, a sparse typed column per key, a presence map to tie them together — and schemaless data gets the same first-class columnar treatment as any declared schema.