Running log of what has been observed, what has been confirmed, and what
is still open. The specification in SPEC.md should never contain a claim
that is not backed by an entry here.
- Source:
../QVD-Sources/downloads/, ~1,089.qvdfiles. - 6 of those are Git LFS pointer files (skipped).
- 36 carry a
.qvdextension but are actually CSV or QVS scripts (skipped: they do not start with<?xmlor<QvdTableHeader). - 1,047 files parse as real QVD headers. 1,043 use CRLF line endings
and a
\r\n\x00header terminator; 4 use LF and\n\x00. QvBuildNospans the 11000s and 50000s. No structural differences have been observed across builds for the features in Stage 1 and 2.
- Terminator
</QvdTableHeader>\r\n\x00or</QvdTableHeader>\n\x00. <Compression>is present and empty in every sample.- Root
Offset+Lengthdescribe the row block inside the body. - Per-field
Offset+Lengthdescribe the symbol table inside the body. - Symbol tables and the row block do not overlap in any observed file.
Type byte histogram across 1,045 files (~25 million entries):
| byte | meaning | count |
|---|---|---|
0x01 |
int32 LE | 3,361,575 |
0x02 |
float64 LE | 922,225 |
0x04 |
nul-terminated UTF-8 string | 10,633,806 |
0x05 |
int32 LE + nul-terminated UTF-8 string | 10,252,234 |
0x06 |
float64 LE + nul-terminated UTF-8 string | 453,558 |
| other | unknown | ~7 (one damaged file) |
All ~25M entries fit this table cleanly.
Example round-trip, file
korolmi/qvdfile/qvdfile/data/tab1.qvd field ID:
- Symbols at
Offset=0, Length=40:06 48 e1 7a 14 ae c7 5e 40 31 32 33 2e 31 32 00-> dual float64123.12, string"123.12"05 7c 00 00 00 31 32 34 00-> dual int32 124, string"124"05 fe ff ff ff 2d 32 00-> dual int32 -2, string"-2"05 01 00 00 00 31 00-> dual int32 1, string"1"
Five rows, BitWidth=3, Bias=-2:
| byte | bits[0..3) | stored | index = stored + Bias | value |
|---|---|---|---|---|
02 |
010 | 2 | 0 | 123.12 |
0b |
011 | 3 | 1 | 124 |
14 |
100 | 4 | 2 | -2 |
1d |
101 | 5 | 3 | 1 |
20 |
000 | 0 | -2 | NULL |
Confirms the stored + Bias rule and the negative-index-means-NULL rule.
Example ProductGroup.qvd: RecordByteSize=2, 17 rows, two fields at
(BitOffset=0, BitWidth=8) and (BitOffset=8, BitWidth=8). Row bytes
00 00 01 01 02 02 ... 10 10 decode with LSB-first byte ordering and
produce identity mappings (0->0, 1->1, ..., 16->16) in each field, as
expected.
Non-byte-aligned widths are common: 816 of 1,045 files have at least one field whose width is not a multiple of 8. Widths of 1..8 bits are the most common; rarer packings (ex. 2+6, 3+5, 1+7) also appear. The examination script confirms that reading each field as a little-endian bit-field starting at the documented offset reproduces sensible indices that then resolve against the symbol table.
- What is the exact interpretation of
NumberFormat/Typevalues likeFIX,MONEY,INTERVAL? Does it affect how dual floats should be displayed? - Are there
Tagscombinations that affect physical representation or only semantics? - Is there any alignment padding inside the symbol block? Currently the examination script successfully reads every symbol entry immediately adjacent to the previous one; no padding has been observed.
- Do any QVDs in the wild carry a non-empty
<Compression>element? So far no samples exist. - Are there multi-table QVD files anywhere? None observed.
Hankiiee/CodeCoverage/.../mynewfile3.qvd: ~7 symbol entries begin with unknown type bytes including0x35,0xaf,0xb5,0xb8,0xc2,0x30. File appears to be corrupted or truncated. Needs a closer look, but should not affect the spec.
A pure-Python decoder (re/decode.py) implemented strictly from
SPEC.md was run over the full corpus with --max-rows 100. Result:
- 1,044 files decoded cleanly.
- 3 files failed, all expected corruption:
ptarmiganlabs/ctrl-q-qvd-viewer/test-data/misc/damaged.qvd(intentionally corrupted test fixture, filename says so).MuellerConstantin/qvd4js/__tests__/data/damaged.qvd(same).Hankiiee/.../mynewfile3.qvd(previously flagged).
The decoder enforces every rule in the spec and fails loudly on any deviation: bit-field overlap, out-of-range symbol indices, trailing bytes in a symbol table, mis-sized row block, unknown type byte, or truncated string payload. None of these triggered on a non-damaged file.
This is strong evidence that Stages 1..3 of the spec are complete and correct for the uncompressed, single-table QVD files present in the public corpus.
A Rust crate in crates/openqvd/ implements the spec independently of
the Python decoder. It exposes Qvd::from_path / Qvd::from_bytes, a
typed Value enum covering the five symbol kinds (int, float, string,
dual-int, dual-float), and a row iterator that yields Vec<Option<Value>>
with explicit NULL handling for bias-based nulls.
Validation: the validate_all example parses 1,044 of 1,047 files in
the corpus; the 3 failures are the same damaged files that failed the
Python decoder. Unit test minimal.rs round-trips a hand-authored
header+symbols+rows payload built from the spec alone.
Known limitations in this first pass:
- Rows block larger than 16 bytes per record uses a slower fallback path; still correct, but could be faster.
TagsandNumberFormatare preserved but not interpreted.- No writer yet.
The writer (crates/openqvd/src/writer.rs) implements SPEC section 7.
Per-column plan:
- Symbols are deduplicated by value and ordered by first occurrence.
Float keys use
to_bits()so-0.0and+0.0are distinct and NaN-distinct bit patterns survive round-trip. - Bias is
-2when any NULL is present, else0. The smallestBitWidththat fits(n_symbols - 1) + (-bias)is chosen. A column with one symbol and no NULLs collapses toBitWidth=0. - Bit offsets are assigned in declaration order; symbol-block offsets
likewise; the record byte size is
ceil(total_bits / 8).
Row packing uses a u128 fast path when the record fits, and a per-field byte-level bit writer otherwise. The XML header emits only the minimal element set from SPEC 7.3 and is byte-deterministic for a given input.
The round_trip example walks a directory, for each QVD runs
read -> write -> read, and compares cells one at a time. A size cap
keeps memory bounded (the current to_write_table materialises all
cells, which amplifies memory for multi-million-row files; future
improvement would stream directly from reader symbol tables).
Result over /workspaces/Sigilweaver/QVD-Sources/downloads with a
128 MiB cap:
total=1145 ok=1093 skipped=49 oversize=0 fail=3
- 1093 valid files round-trip with cell-for-cell equivalence.
- 3 failures are the same damaged fixtures that already fail the reader.
- 49 skipped are non-QVD bytes (LFS pointers, misnamed CSV/QVS).
Round-trip parity also holds on smaller caps (8 MiB: 1053/1053), so the writer is not sensitive to corpus subsets.
A small openqvd binary wraps the library with five subcommands:
stat, head, csv, json, rewrite. This is the first
user-facing surface and makes the project useful as a standalone tool,
not just a library.
- Stream
to_write_tableto remove the memory amplification for very large files (>128 MiB). None exist in the public corpus today.
crates/openqvd-py is a mixed-layout maturin crate that exposes the Rust
library to Python via PyO3 0.28 and pyo3-arrow 0.17. The native extension
ships as a single .so file; the pure-Python wrapper in
python/openqvd/ provides read(), schema(), and write() plus
Polars integration (read_qvd, scan_qvd, df.qvd.write,
lf.qvd.collect_and_write).
to_record_batch infers Arrow data types from the QVD
NumberFormat/Type hint first, then falls back to symbol value types.
DATE → Date32 (shifted by 25 569 days to the Unix epoch), TIMESTAMP →
Timestamp(µs), TIME → Duration(µs) (fractional days × 86 400 000 000).
Int/DualInt → Int64, Float/DualFloat → Float64, String → LargeUtf8,
empty symbol table → Null.
Filters are resolved against symbol tables before row iteration.
For each filter, resolve_filter() computes a set of symbol indices
that satisfy the predicate (or, for NOT_IN, the set of indices that
fail). During row unpacking, rows whose symbol indices do not match
any resolved filter are skipped entirely - no value decoding or Arrow
append occurs for those rows.
Supported operators: eq, is_in, not_in, is_null, is_not_null.
All comparisons are string-based (against value_as_str of each symbol).
Full corpus run (1 089 files, 42 LFS stubs skipped):
- 1 044 OK - matches the Rust reader baseline exactly.
- 3 FAIL - same damaged fixtures that fail the Rust reader.
34 pytest tests cover:
- Core API: read, schema, projection, metadata, row counts (9 tests)
- Predicate pushdown: eq, is_in, not_in, combined, no-match, error handling (7 tests)
- Write round-trip: basic, with table name, value preservation, NULLs (5 tests)
- Polars: read, scan, projection, filters, write, lazy write (9 tests)
- Pandas: to_pandas, values, projection, NULLs (4 tests)