# Schemaless, Not Structureless: A True Columnar Index for Log Data

## The problem: logs have no fixed shape

Logs are the wild west of telemetry. Every service logs differently, every team adds its own fields, and enrichment pipelines bolt on more — pod names, regions, build versions — long after the code was written. Each log line arrives as a loose bag of facets — key-value attributes — and no two lines carry quite the same bag. For decades, storage engines have forced an uncomfortable choice on this data: keep it flexible and query it slowly, or make it fast and lose the flexibility.

We built a storage index that removes that tradeoff: a **schemaless index** that gives every facet — whether it appears on a billion log lines or just three — a true, typed, compressed column of its own, created automatically at ingest time. This post explains how it works, why it is fundamentally faster than JSON-blob and row-oriented designs, and why it can keep absorbing new fields without ever asking for a schema migration. The design applies to any telemetry — traces, events, user sessions — but logs are where schema chaos is most extreme, so logs are the lens we'll use.

## The problem: logs have no fixed shape

Consider four log lines from a single application fleet:

Three things make this data hard to store well:

- **The key space is unbounded.** Across a fleet of services, teams, and deploys, the set of distinct facets runs into the thousands — and nobody can enumerate it up front.
- **Each line is sparse.** Any single log line carries only a handful of those thousands of keys. Most keys are absent from most lines.
- **The shape changes constantly.** Tomorrow's deploy starts logging `checkout.retries`; next week's adds `feature.flag.variant`. New keys appear daily, without warning or coordination.

## The usual answers, and what they cost

Three architectural patterns dominate how analytical systems store log facets today.

**1\. The JSON blob.** Store each line's facets as one serialized JSON string in a single column. Ingest is trivial and nothing is ever dropped — but every query pays the full price. Asking "what is the average `duration_ms` for the payments service?" means reading _every_ blob, parsing _every_ blob, and extracting one field from each, even though the other 99% of each blob is irrelevant to the question. Compression suffers too.

**2\. Row-oriented document storage.** Store each log line as a document, retrieved whole. This is excellent for "show me this one log line" — and structurally wrong for analytics. Faceting on a single key across ten million lines means reading ten million whole documents. Such systems stay fast only by building inverted indexes on _everything_ in advance, multiplying ingest cost and storage for facets that may never be queried.

**3\. The rigid columnar schema.** Classic columnar storage delivers exactly the read pattern log analytics wants: scan only the columns the query touches. But it demands the schema up front. Every new log field is a schema migration; every undeclared field is either dropped or stuffed into… a JSON blob, returning us to pattern 1.

|  | JSON blob | Row/document store | Rigid columnar | **Schemaless index** |
| --- | --- | --- | --- | --- |
| Accepts any shape | ✅ | ✅ | ❌ | ✅ |
| Reads only what the query needs | ❌ | ❌ | ✅ | ✅ |
| Columnar compression | ❌ | ❌ | ✅ | ✅ |
| New key requires migration | No | No | **Yes** | No |
| Values stored typed (no re-parsing) | ❌ | partly | ✅ | ✅ |

The fourth column is the subject of the rest of this post.

## The schemaless index: every key becomes a real column

The core idea is simple to state: **decompose log lines at ingest time, so that every distinct facet gets its own typed, sparse column — automatically.** Flexibility is preserved at the boundary where data arrives; full columnar structure exists everywhere data is stored and read.

Four structures make this work:

**A key dictionary.** Every distinct key, together with its inferred value type, gets one entry in a compact dictionary — the index's table of contents.

**A sparse column per key.** Each key's values are stored contiguously, in arrival order — but _only for the lines that actually have the key_.

**A presence map per key.** A compact bitmap records exactly which lines carry the key. This is the bridge between "log line #7" and "the 3rd value in this column".

**Typed values.** Types are inferred per key at ingest: numbers are stored as numbers, strings as strings, booleans as booleans.

## Why this is fast

Each speed advantage falls directly out of a structure above — no benchmarks required to see why.

**I/O proportional to the question, not the data.** A query touching one key out of 200 reads roughly 1/200th of the facet data. The blob design reads and parses **100%** of it to answer the same question.

**No parsing, ever.** Values are stored already decoded and typed. Reading an integer is reading an integer — not tokenizing a JSON string.

**Compression that actually works.** A column of 10,000 status codes is highly repetitive — exactly what block compressors are built for.

**Encodings tailored to each key.** Per-key columns unlock per-key encoding choices.

**Sparsity is free.** Rare keys cost almost nothing — a key on 1% of lines uses ~1% of the storage of a dense column, plus a tiny presence map.

## Built to grow

Extensibility is not a feature added on top of this design — it _is_ the design.

**The data model evolves itself.** New keys become columns automatically; new value types coexist as sibling columns.

**Storage decisions are made per key.** Compression codec and presence-map strategy are chosen per column.

**The key is the unit of optimization.** Because every facet already lives in its own column, future accelerations attach naturally at the same granularity.

## Conclusion

The flexibility-versus-performance tradeoff for log data was never a law of nature. It was an artifact of one decision: treating a flexible log line as an opaque blob. Move the decomposition to ingest — a dictionary of keys, a sparse typed column per key, a presence map to tie them together — and schemaless data gets the same first-class columnar treatment as any declared schema.
