rumi
- Specification 0.1.0
- Status Draft
- Date 2026-09-24
- License GPLv3
rumi is stateless raster storage for AI4EO. Its GeoTIFF-inspired format stores
compressed raster tensors with up to four dimensions. A file contains either an
Image (B, Y, X) or a temporal Cube (T, B, Y, X).
The format has the following properties:
- Stateless reads. The external binary header locates every compressed frame
without parsing the
.rumifile or retaining state between requests. - OpenZL compression. Stores each frame as an independent, self-contained OpenZL frame.
- Predictable layout. Files with the same band and frame counts begin their frame data at the same byte.
- Described bands and time. Names every band and labels every time step in a trailer after the frame data.
- Canonical structure. Restricts the file to one fixed IFD, one frame order, and contiguous frame data.
- Fixed-size georeferencing. Stores an EPSG CRS and affine transform in a 160-byte block derived from GeoTIFF tags, with a defined value for ungeoreferenced rasters.
The filename extension for the format is .rumi.
The key words MUST, MUST NOT, SHOULD, and MAY are to be interpreted as described in RFC 2119 when, and only when, they appear in capitals.
Scope
This document defines:
- the rumi data model and frame layouts;
- the GeoTIFF-inspired rumi file profile;
- the band descriptions and time coordinates every file carries; and
- the binary layout of the external rumi header blob.
It does not define the OpenZL frame format or a catalogue format.
This document contains everything needed to implement a rumi reader. The only supported writer is the one provided by rumi; independent writers are outside the compatibility policy.
Unless a section says otherwise, all integer arithmetic used to validate or derive sizes, counts, and offsets is exact. A reader MUST reject an input when a required result cannot be represented by its implementation.
All multi-byte numeric values defined by rumi are little-endian. This includes the file header, directory entries and values, decoded sample components, the trailer, and the external header blob. OpenZL defines the bytes inside a compressed frame; after decoding, each multi-byte sample component is little-endian. An API may convert decoded samples to the host's native byte order.
Selective read model
This section is informative. A selective read combines band and time positions,
a spatial window, the external header that belongs to the file, and access to
the .rumi file.
- The window identifies the tile rows and columns it intersects.
frame_unitdetermines which band and time positions are part of each frame index and which are decoded inside a frame.- The header's packed byte counts reconstruct the offset and length of every required frame.
- Each selected range is decoded as an independent OpenZL frame and its requested samples are placed in the output array.
For example, consider a Cube with shape (T=3, B=4, Y=1024, X=1024), nominal
tiles of 256 × 256, and frame_unit = 0. Its spatial grid is 4 × 4, so
g = 16 and N = g * B * T = 192.
A request for time position t = 2, bands b = 1 and b = 3, and the window
y = [512, 768), x = [256, 512) covers the tile at row = 2, col = 1.
spatial = row * tiles_across + col
= 2 * 4 + 1
= 9
index(b=1, t=2) = (9 * 4 + 1) * 3 + 2 = 113
index(b=3, t=2) = (9 * 4 + 3) * 3 + 2 = 119
The reader reconstructs the offsets of frames 113 and 119 as defined in
Offset reconstruction, fetches those two ranges, and
leaves the other 190 frames untouched.
Data model
rumi uses the following data model.
| term | shape | definition |
|---|---|---|
| tile | (h, w) |
one band and time step at one tile location |
| cell | (B, h, w) or (T, B, h, w) |
all samples at one tile location |
| Image | (B, Y, X) |
one raster grid; a rumi file with T = 1 |
| Cube | (T, B, Y, X) |
T time-ordered, grid-aligned Images in one file |
| ImageCollection | — | a set of Images that need not share a grid |
| CubeCollection | — | a set of Cubes that need not share a grid |
T, B, Y, and X denote time count, band count, image length, and image
width. h and w denote the actual dimensions of a tile; they may be smaller
than the nominal tile dimensions at the image boundary.
The bands and time steps of an Image or Cube are described in the Trailer.
Collections are represented outside the file, for example by a catalogue of rumi files. Their representation is out of scope.
Frames
tile and cell belong to the logical raster model. A frame is the physical
unit of compression and random access. Each frame is a self-contained OpenZL
frame that carries its own graph and codec parameters; neither is stored in the
IFD or external header. Its decoded shape and sample order are specified by
frame_unit.
A frame MUST contain exactly one of:
- one tile for one
(b, t)pair at one tile location; or - one cell at one tile location.
frame_unit
frame_unit selects one of the decoded layouts below. b, t, h, and w
mean band, time, height, and width. The rightmost axis changes fastest.
| frame_unit | decoded frame | valid when |
|---|---|---|
0 |
h w |
any B and T |
1 |
b h w or t h w |
exactly one of B, T exceeds 1 |
2 |
h w b or h w t |
exactly one of B, T exceeds 1 |
3 |
b t h w |
B > 1 and T > 1 |
4 |
t b h w |
B > 1 and T > 1 |
5 |
b h w t |
B > 1 and T > 1 |
6 |
t h w b |
B > 1 and T > 1 |
7 |
h w b t |
B > 1 and T > 1 |
8 |
h w t b |
B > 1 and T > 1 |
9 |
h w |
B > 1 and T > 1 |
Units 1 and 2 place one non-spatial axis around h w. That axis is b when
B > 1, and t when T > 1.
Units 0 and 9 differ only in the order the index walks the two axes at one
tile location: b then t for 0, t then b for 9. That order is
observable only when both B and T exceed 1, which is why 9 is valid
nowhere else.
The diagram shows how the two frame types use the band and time axes. For a
tile, b and t select the frame. For a cell, they are part of the decoded
frame.
Within a decoded frame, h w MUST stay together and in that order.
The registry is complete and append-only; existing values MUST NOT be
reassigned. A reader MUST reject any frame_unit value or (B, T) combination
not listed above. A valid frame_unit determines the decoded sample order and
the number of frames N; it does not otherwise change the file structure.
Choosing a frame unit
This section is informative. Units 0 and 9 provide the finest access: one
frame contains one tile for one band and one time step. Every other unit places
one cell in each frame, allowing OpenZL to model correlation between bands,
time steps, or both.
A frame is the smallest unit a reader decodes, so the order of samples inside
it changes only compression. A frame is a tile or a whole cell, never an
intermediate group, and the index order of units 0 and 9 decides which
tiles sit next to each other.
An axis before h w is stored as contiguous planes. An axis after h w is
interleaved within each pixel. When both axes precede h w, the axis next to
h w varies between adjacent planes.
The best unit depends on the expected reads and the data. Compression SHOULD be measured when more than one unit fits the access pattern.
Frame index
row and col select a tile location in the spatial grid. Tile locations are
traversed in row-major order.
tiles_across = ceil(image_width / tile_width)
tiles_down = ceil(image_length / tile_length)
g = tiles_across * tiles_down
When a frame holds a cell, there is one frame per tile location.
frame_index(row, col) = row * tiles_across + col
N = g
When a frame holds a tile, the index also walks the band and time axes.
spatial = row * tiles_across + col
frame_unit 0: frame_index = (spatial * B + b) * T + t
frame_unit 9: frame_index = (spatial * T + t) * B + b
N = g * B * T
Validating frame_unit
Tag 65000 stores frame_unit in the file. Its value MUST be registered and
valid for B and T.
The entry counts of TileOffsets and TileByteCounts MUST also match the frame
unit:
count == g * B * T when the frame holds a tile
count == g when the frame holds a cell
A reader MUST reject a file when either count is wrong. The counts do not
distinguish 0 from 9, or one full-frame layout from another; tag 65000
does.
Frame order
Frames MUST be stored in increasing frame_index order. TileOffsets and
TileByteCounts MUST use the same order.
This order allows the external header to reconstruct offsets with a prefix sum.
A reader MUST reject a file whose TileOffsets do not match the reconstructed
offsets in frame-index order.
Sample encodings
sample_format gives the sample type and bits_per_sample its width in bits.
Unsigned integers use ordinary binary representation, and signed integers use
two's-complement representation. IEEE formats use the IEEE 754 binary16,
binary32, or binary64 encoding named in the table. Boolean samples use one
decoded byte whose value is 0 or 1, although their logical width is one bit.
A complex sample stores two equal-width components: real first, then imaginary.
For complex formats, bits_per_sample is their combined width.
| sample_format | meaning |
|---|---|
1 |
unsigned integer |
2 |
signed integer |
3 |
IEEE floating point |
6 |
complex IEEE floating point |
100..103, 107, 108 |
rumi-private ML floating point types |
Only the following pairs are valid:
| sample_format | bits_per_sample | decoded bytes | encoding |
|---|---|---|---|
| 1 | 1 | 1 | boolean |
| 1 | 8 | 1 | unsigned 8-bit integer |
| 1 | 16 | 2 | unsigned 16-bit integer |
| 1 | 32 | 4 | unsigned 32-bit integer |
| 1 | 64 | 8 | unsigned 64-bit integer |
| 2 | 8 | 1 | signed 8-bit integer |
| 2 | 16 | 2 | signed 16-bit integer |
| 2 | 32 | 4 | signed 32-bit integer |
| 2 | 64 | 8 | signed 64-bit integer |
| 3 | 16 | 2 | IEEE 16-bit floating point |
| 3 | 32 | 4 | IEEE 32-bit floating point |
| 3 | 64 | 8 | IEEE 64-bit floating point |
| 6 | 32 | 4 | complex IEEE floating point, 16-bit components |
| 6 | 64 | 8 | complex IEEE floating point, 32-bit components |
| 6 | 128 | 16 | complex IEEE floating point, 64-bit components |
| 100 | 8 | 1 | float8 E4M3FN |
| 101 | 8 | 1 | float8 E5M2 |
| 102 | 16 | 2 | bfloat16 |
| 103 | 8 | 1 | float8 E8M0FNU |
| 107 | 8 | 1 | float8 E4M3FNUZ |
| 108 | 8 | 1 | float8 E5M2FNUZ |
A reader MUST reject any pair not listed above.
The pairs (1, 2), (1, 4), (2, 2), (2, 4), (5, 32), (5, 64),
(104, 6), (105, 6), and (106, 4) are reserved. The sample_format
values 5, 104, 105, and 106 are also reserved and MUST NOT be assigned
another meaning.
bfloat16 has one sign bit, eight exponent bits, and seven fraction bits, with
the exponent and special values of IEEE binary32. E4M3FN, E4M3FNUZ,
E5M2, and E5M2FNUZ use the
ONNX float8 encodings.
E8M0FNU uses the corresponding encoding in
the OCP Microscaling Formats (MX) Specification
1.0.
All bands and time steps in a file MUST use the same pair.
bits_per_sample is the logical width of a sample, not necessarily its storage
stride. Boolean samples occupy one decoded byte and every byte MUST be 0 or
1. Packed boolean storage MUST NOT be used.
The decoded frame size is defined by:
decoded_samples = h * w when the frame holds a tile
B * T * h * w when the frame holds a cell
bytes_per_sample = the decoded bytes in the sample-encoding table
decoded_frame_bytes = decoded_samples * bytes_per_sample
An absent axis contributes a factor of one. Resource limits
applies to decoded_frame_bytes.
DLPack representation
An implementation that exports decoded samples through DLPack MUST use the
following (code, bits, lanes) values. Every tensor is native-endian CPU memory
and one tensor element corresponds to one rumi sample. The float8 codes require
DLPack 1.1 or newer.
| encoding | DLPack (code, bits, lanes) |
|---|---|
| signed integers | (kDLInt=0, bits_per_sample, 1) |
| unsigned integers | (kDLUInt=1, bits_per_sample, 1) |
| boolean | (kDLBool=6, 8, 1) |
| IEEE floats | (kDLFloat=2, bits_per_sample, 1) |
| complex IEEE floats | (kDLComplex=5, bits_per_sample, 1) |
| bfloat16 | (kDLBfloat=4, 16, 1) |
| float8 E4M3FN | (kDLFloat8_e4m3fn=10, 8, 1) |
| float8 E4M3FNUZ | (kDLFloat8_e4m3fnuz=11, 8, 1) |
| float8 E5M2 | (kDLFloat8_e5m2=12, 8, 1) |
| float8 E5M2FNUZ | (kDLFloat8_e5m2fnuz=13, 8, 1) |
| float8 E8M0FNU | (kDLFloat8_e8m0fnu=14, 8, 1) |
The file's logical width and DLPack's element width differ for boolean data; the decoded storage width connects them. A reader MUST NOT silently cast, reinterpret, or widen a sample to satisfy a consumer.
Bit-packed arrays
Time residuals and frame byte-count residuals use the same bit packing.
Values are stored consecutively using bits bits each. Bit j of value i
occupies bit position i * bits + j, with bits numbered from the least
significant bit of each byte. For n values, the region is exactly
ceil(n * bits / 8) bytes. Unused bits in the final byte MUST be zero.
When bits is 0, the region is empty and every value is zero.
File profile
A file is rumi compliant when all of the following hold.
- It begins with the rumi file header and contains exactly one rumi IFD.
- It is tiled and has no overviews, masks, strips, or auxiliary IFDs.
- Its IFD precedes the frame data, and its tags, values, and frame placement follow Fixed IFD, without gaps or padding.
- Its sample encoding is listed in Sample encodings.
- Its georeferencing follows Georeferencing.
- Each frame is a self-contained OpenZL frame.
FrameUnit,TileOffsets, andTileByteCountssatisfy Validating frame_unit.- Every frame is present, every byte count is greater than zero, and the frames form one contiguous run in frame-index order.
- It ends with the trailer defined in Trailer.
File header
Every rumi file begins with this 16-byte header.
| offset | size | type | name |
|---|---|---|---|
| 0 | 4 | bytes | magic |
| 4 | 2 | uint16 | version |
| 6 | 2 | uint16 | reserved |
| 8 | 8 | uint64 | ifd_offset |
The magic bytes spell ASCII RUMI: 52 55 4D 49. The current version is 1,
reserved is zero, and ifd_offset is 16. A reader MUST reject any other
value.
The IFD uses the 20-byte entry layout and tag numbers derived from BigTIFF, but rumi defines its own tags, placement, and alignment. A rumi file is not a TIFF, BigTIFF, or GeoTIFF file.
Fixed IFD
A rumi IFD MUST contain exactly the tags below, in rising tag order. A reader that validates the file or builds an external header from it MUST reject any other tag.
B is samples_per_pixel, T is time_count, and N is the frame count.
The IFD begins with the eight-byte entry count 13 and ends with an eight-byte
zero offset for the next IFD. Each entry has this layout:
| offset | size | type | name |
|---|---|---|---|
| 0 | 2 | uint16 | tag |
| 2 | 2 | uint16 | type |
| 4 | 8 | uint64 | count |
| 12 | 8 | bytes | value or offset |
The type codes are 3 for SHORT, 4 for LONG, 12 for DOUBLE, and 16
for LONG8. These represent uint16, uint32, IEEE 754 binary64, and uint64,
respectively.
| tag | name | type | count |
|---|---|---|---|
| 256 | ImageWidth | LONG | 1 |
| 257 | ImageLength | LONG | 1 |
| 258 | BitsPerSample | SHORT | B |
| 277 | SamplesPerPixel | SHORT | 1 |
| 322 | TileWidth | SHORT | 1 |
| 323 | TileLength | SHORT | 1 |
| 324 | TileOffsets | LONG8 | N |
| 325 | TileByteCounts | LONG | N |
| 339 | SampleFormat | SHORT | B |
| 34264 | ModelTransformationTag | DOUBLE | 16 |
| 34735 | GeoKeyDirectoryTag | SHORT | 16 |
| 65000 | FrameUnit | SHORT | 1 |
| 65001 | TimeCount | LONG | 1 |
Every file carries all 13 tags.
ImageWidth, ImageLength, TimeCount, TileWidth, TileLength, and
SamplesPerPixel MUST be greater than zero.
BitsPerSample and SampleFormat MUST contain B repetitions of one pair
listed in Sample encodings.
Placement
The IFD starts at byte 16, immediately after the rumi file header. Its size is
8 + 20 * 13 + 8 = 276 bytes: an eight-byte entry count, 13 entries, and an
eight-byte zero offset for the next IFD.
Values of 8 bytes or less MUST be stored in the IFD entry. Larger values MUST
follow the IFD in rising tag order, without gaps. Unused bytes in an inline
value MUST be zero. For an external value, the last eight bytes of the entry
store its uint64 file offset.
The frame data starts immediately after the last external value. Padding or alignment bytes MUST NOT be inserted in the external area.
Deriving base_frame_offset
The IFD size is fixed. Only values larger than 8 bytes contribute to the external area before the frames.
external = (2 * B if B >= 5 else 0) # 258 BitsPerSample
+ (8 * N if N >= 2 else 0) # 324 TileOffsets
+ (4 * N if N >= 3 else 0) # 325 TileByteCounts
+ (2 * B if B >= 5 else 0) # 339 SampleFormat
+ 128 # 34264 ModelTransformationTag
+ 32 # 34735 GeoKeyDirectoryTag
base_frame_offset = 16 + 276 + external
= 292 + external
The result MUST match the first entry of TileOffsets. Files with the same B
and N start their frame data at the same byte.
Georeferencing
Every rumi file carries ModelTransformationTag and GeoKeyDirectoryTag.
ModelPixelScaleTag (33550), ModelTiepointTag (33922),
GeoDoubleParamsTag (34736), and GeoAsciiParamsTag (34737) MUST NOT appear.
Files without georeferencing use the values in
Undefined georeferencing.
ModelTransformationTag
ModelTransformationTag stores a 4 × 4 matrix as 16 doubles in row-major order.
North-up rasters set the rotation terms to zero.
Given affine coefficients
(x_res, row_rot, x_origin, col_rot, y_res, y_origin), the matrix is
x_res row_rot 0 x_origin
col_rot y_res 0 y_origin
0 0 0 0
0 0 0 1
The third row MUST be zero.
GeoKeyDirectoryTag
rumi represents a CRS by EPSG code. GeoKeyDirectoryTag contains a four-short
header followed by three four-short keys, for a total of 16 shorts.
| key | id | value |
|---|---|---|
| GTModelTypeGeoKey | 1024 | 1 projected or 2 geographic |
| GTRasterTypeGeoKey | 1025 | 1 PixelIsArea or 2 PixelIsPoint |
| GeographicTypeGeoKey or ProjectedCSTypeGeoKey | 2048 or 3072 | EPSG code |
The four-short header MUST be (1, 1, 0, 3). Each key is stored as
(key_id, 0, 1, value), in the order shown above.
A geographic CRS uses key 2048; a projected CRS uses key 3072. The key MUST
agree with GTModelTypeGeoKey. The EPSG code MUST be defined by the EPSG
registry and fall between 1024 and 32766.
A writer MUST obtain the CRS type from the EPSG registry or a PROJ database. It
MUST NOT infer the type from the numeric code. A reader obtains the type from
GTModelTypeGeoKey.
Undefined georeferencing
An ungeoreferenced file still carries both georeferencing tags.
ModelTransformationTag uses the following matrix.
1 0 0 0
0 1 0 0
0 0 0 0
0 0 0 1
GeoKeyDirectoryTag uses the same header and key representation:
| key | id | value |
|---|---|---|
| GTModelTypeGeoKey | 1024 | 0 |
| GTRasterTypeGeoKey | 1025 | 1 or 2 |
| GeographicTypeGeoKey | 2048 | 0 |
A reader MUST interpret GTModelTypeGeoKey = 0 as no CRS and MUST ignore the
transformation matrix.
The two tags always occupy 160 bytes.
No other CRS representation is permitted. This excludes WKT, PROJ strings, ESRI codes, user-defined CRS values, engineering, compound and vertical CRSs, and coordinate epochs.
Trailer
Every rumi file ends with one trailer. It begins immediately after the last frame, and the file ends immediately after it. The trailer describes every band and labels every time step.
The trailer follows the frame data so that its size never moves a frame. A reader MUST NOT use it to locate, decode, or convert frames. The external header contains no band descriptions or time coordinates.
+----------------+---------------+-------------+-------------------------------+
| magic, version | band_texts[B] | time fields | time_residuals[C] |
+----------------+---------------+-------------+-------------------------------+
6 bytes S bytes 22 bytes ceil(C * time_bits / 8) bytes
S is the size of the band texts, defined in
Band descriptions. The trailer size MUST be exactly
28 + S + ceil(C * time_bits / 8) bytes and MUST contain no padding. All
multi-byte fields are little-endian.
Magic and version
| offset | size | type | name |
|---|---|---|---|
| 0 | 4 | uint32 | magic |
| 4 | 2 | uint16 | version |
The four magic bytes spell ASCII TAIL: 54 41 49 4C. Read as a little-endian
uint32, they equal 0x4C494154. A reader MUST reject any other value.
The current trailer version is 1. A reader that implements version 1 MUST
reject any other value.
Band descriptions
The version is followed by one text for each of the B bands, in band order.
Each text is stored as its length in bytes followed by the bytes themselves.
| size | type | name |
|---|---|---|
| 2 | uint16 | n[b] |
| n[b] | bytes | text |
S = sum(2 + n[b]) for 0 <= b < B
Each text MUST be valid UTF-8, at least one byte long, and free of the byte
0x00. Two texts in one file MUST NOT be equal as byte sequences. A reader MUST
reject a trailer whose texts break these rules.
This paragraph is informative. The recommended text gives the band name, a short description, and the wavelength, as in this text for the Sentinel-2 red band:
B4, Red, 664.5nm (S2A) / 665nm (S2B)
Time coordinates
The time fields follow the last band text. Every time step has a coordinate. An Image without a single acquisition instant, such as a DEM or an annual composite, uses an interval.
Let C be the number of coordinates stored in the trailer:
C = T when time_type is 2 (instant)
C = 2 * T when time_type is 1 (interval)
For instants, time(i) is the coordinate of time step i.
For intervals, time step i covers [time(2i), time(2i + 1)). Each step stores
its own start and end, so intervals may leave gaps.
Offsets are relative to the first byte after the last band text.
| offset | size | type | name |
|---|---|---|---|
| 0 | 1 | uint8 | time_type |
| 1 | 1 | uint8 | time_bits |
| 2 | 8 | int64 | time_epoch |
| 10 | 8 | int64 | time_step |
| 18 | 4 | uint32 | time_scale |
time_type
| value | meaning |
|---|---|
1 |
interval; each step is bounded by two coordinates |
2 |
instant; the step is a point in time |
A reader MUST reject any other value.
Instant coordinates MUST be non-decreasing.
Interval coordinates MUST satisfy:
time(2i) < time(2i + 1) for 0 <= i < T
time(2i + 1) <= time(2i + 2) for 0 <= i < T - 1
Intervals may meet or leave gaps, but MUST NOT overlap.
time_epoch, time_step and time_scale
time_scale is the number of seconds represented by one coordinate unit. It
MUST be 86400 when every instant or interval endpoint is an exact whole-day
offset from 1970-01-01T00:00:00Z; otherwise it MUST be 1. A reader MUST
reject any other value.
Coordinate time(i) is a signed offset of time(i) * time_scale seconds from
1970-01-01T00:00:00Z; negative values represent times before that epoch. rumi
follows POSIX time and does not represent leap seconds.
time_epoch MUST equal time(0). time_step is the slope of the prediction
line used to encode the remaining coordinates:
time_step = 0 if C < 2
time_step = round((time(C-1) - time(0)) / (C-1)) otherwise
round chooses the nearest integer; exact halves round toward positive
infinity.
time_bits
The number of bits used to encode each residual, as defined in
Time residuals. It MUST be between 0 and 64.
Time residuals
After the time fields, the trailer stores the C coordinates as residuals
against a straight line through time_epoch and time_step.
An axis whose coordinates lie on this line requires no packed region. Otherwise, the trailer stores their deviations from the line.
Encoding
Let time(i) be coordinate i in time_scale units.
predicted(i) = time_epoch + i * time_step
residual(i) = time(i) - predicted(i)
A residual MUST fit in int64. It is mapped to an unsigned integer with zigzag
encoding:
zigzag(x) = 2 * x if x >= 0
-2 * x - 1 otherwise
packed(i) = zigzag(residual(i))
time_bits = bit_length(max(packed))
bit_length(0) is 0.
A writer MUST use the minimum time_bits that represents the largest packed
residual. residual(0) is always zero.
A reader MUST reject a trailer unless time_scale satisfies the rule above and
the decoded coordinates reproduce the recorded time_epoch, time_step, and
minimum time_bits. Every decoded coordinate MUST fit in int64 and satisfy
the ordering required by time_type.
Packing
The C packed values are stored as defined in
Bit-packed arrays, at time_bits bits each.
Decoding
residual(i) = packed(i) >> 1 if packed(i) is even
-((packed(i) >> 1) + 1) otherwise
time(i) = time_epoch + i * time_step + residual(i)
Header blob
The header blob is a binary record stored outside the rumi file. It contains the raster fields and frame byte counts needed to locate a frame without parsing the file.
The blob consists of a fixed 32-byte header followed by packed frame byte counts. All multi-byte fields are little-endian.
+---------------+----------------------------------+
| Header | frame_byte_counts[N] |
+---------------+----------------------------------+
32 bytes ceil(N * count_bits / 8) bytes
N is derived as defined in Frame index.
The blob size MUST be exactly 32 + ceil(N * count_bits / 8) bytes and contain
no padding.
The blob stores no offsets. They are reconstructed as defined in Deriving base_frame_offset and Offset reconstruction.
Header fields
| offset | size | type | name |
|---|---|---|---|
| 0 | 4 | uint32 | magic |
| 4 | 2 | uint16 | version |
| 6 | 4 | uint32 | image_width |
| 10 | 4 | uint32 | image_length |
| 14 | 4 | uint32 | time_count |
| 18 | 2 | uint16 | tile_width |
| 20 | 2 | uint16 | tile_length |
| 22 | 2 | uint16 | samples_per_pixel |
| 24 | 1 | uint8 | bits_per_sample |
| 25 | 1 | uint8 | sample_format |
| 26 | 1 | uint8 | frame_unit |
| 27 | 4 | uint32 | count_min |
| 31 | 1 | uint8 | count_bits |
magic
The magic value is 0x45564F4C, represented on the wire as 4C 4F 56 45. A
reader MUST reject any other value.
version
The current binary format version is 1. A reader that implements version 1
MUST reject any other value.
image_width, image_length and time_count
The raster dimensions. Width and length are in pixels and match ImageWidth and
ImageLength. time_count matches the value of TimeCount.
All three MUST be greater than zero. A time_count of 1 is an Image; anything
larger is a Cube.
tile_width and tile_length
The nominal tile dimensions in pixels. They match TileWidth and TileLength.
Both values MUST be greater than zero.
The grid size is defined in Frame index.
Edge tiles are clipped to the image bounds and are not padded. A reader MUST derive their dimensions from the image shape and grid position.
samples_per_pixel
The band count B. This matches SamplesPerPixel in the IFD and MUST be at least 1.
bits_per_sample and sample_format
The sample type and width, as defined in Sample encodings.
frame_unit
The frame layout, as defined in frame_unit. Together with the
image shape it determines N. Tag 65000 carries the same value inside the
file.
count_min and count_bits
The frame byte count encoding, as defined in
Frame byte counts. count_bits MUST be between 0 and
32.
Frame byte counts
After the fixed header, the blob stores the compressed size of each frame in frame-index order. Every count MUST be greater than zero. Each count is encoded as a residual from the minimum count.
Encoding
Let c[i] be the byte count of frame i.
count_min = min(c)
count_bits = 0 if max(c) == count_min else bit_length(max(c) - count_min)
bit_length(x) is the minimum number of bits required to represent x.
A writer MUST use these values. A reader MUST reject a blob if count_min is
not the minimum decoded count or count_bits is not the minimum required width.
A file with one frame always has count_bits = 0.
Packing
The N residuals c[i] - count_min are stored as defined in
Bit-packed arrays, at count_bits bits each.
Decoding
c[i] = count_min + residual[i]
Every reconstructed count MUST fit in uint32.
Creating the header blob
A writer or header builder MUST create the blob from the finalized rumi file. Every duplicated value MUST match:
| header blob | rumi IFD |
|---|---|
image_width |
ImageWidth |
image_length |
ImageLength |
time_count |
TimeCount |
tile_width |
TileWidth |
tile_length |
TileLength |
samples_per_pixel |
SamplesPerPixel |
bits_per_sample |
every BitsPerSample value |
sample_format |
every SampleFormat value |
frame_unit |
FrameUnit |
frame_byte_counts[i] |
TileByteCounts[i] |
frame_byte_counts and TileByteCounts MUST each contain N entries. The
writer or header builder MUST NOT produce the blob if any comparison fails.
This check happens when the blob is created. A stateless reader can then treat
the blob as authoritative and does not need to read or compare the IFD,
TileOffsets, TileByteCounts, the trailer, or the total file size before
reading a frame. How an application keeps a blob associated with its file is
outside the scope of this specification.
Offset reconstruction
The blob stores no frame offsets. A reader derives base_frame_offset and
reconstructs the offsets in frame-index order.
offset[0] = base_frame_offset
offset[idx+1] = offset[idx] + frame_byte_counts[idx]
Every reconstructed offset MUST fit in uint64.
A reader passes frame_byte_counts[idx] bytes at offset[idx] to the OpenZL
decoder.
The byte just past the last frame is where the trailer begins.
trailer_offset = offset[N-1] + frame_byte_counts[N-1]
trailer_offset locates the band descriptions and time coordinates. Reading a
frame does not require parsing the trailer.
The offset of frame k can also be expressed as
offset[k] = base_frame_offset
+ k * count_min
+ sum(residual[0:k])
Resource limits
A reader MUST check derived sizes before allocating memory or decoding a frame.
This includes the band text lengths and the coordinate count C in the trailer.
An operation that exceeds the reader's resource limits MUST fail before the
allocation or decode. This does not make the rumi file invalid.
Changelog
- 0.1.0. Initial draft.