rumi

Working recipes

Examples

Complete examples for writing an Image or Cube, reading only a requested region, moving the same read between local and remote storage, and producing batches for machine-learning frameworks.

Install the reader and writer

Rumi is designed to expose the same reader and writer through Python, R, and Julia. Reading uses the Rumi core; writing also uses GeoZL to build and compress each frame.

python -m pip install rumi-eo
import geozl
import numpy as np
import rumi
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

Shape convention

An Image is (B, Y, X). A temporal Cube is (T, B, Y, X). Selections use zero-based positions, and a window is always (row, column, height, width).

Write and read one Image

This complete round trip stores a four-band Image in 256 × 256 spatial cells. Each frame contains all four bands of one cell.

image = np.random.default_rng(0).integers(
    0, 4096,
    size=(4, 1024, 1024),
    dtype=np.uint16,
)

frames = rumi.frames(
    image,
    "b (row h) (col w) -> row col (b h w)",
    tile_size=256,
)

for frame in frames:
    graph = geozl.graph(frame.data, "planar>zigzag>zstd")
    frame.compressed = geozl.compress(frame.data, graph=graph)

image_path, image_header = rumi.write(
    "image.rumi",
    frames,
    bands=[
        "B2, Blue, 496.6nm (S2A) / 492.1nm (S2B)",
        "B3, Green, 560nm (S2A) / 559nm (S2B)",
        "B4, Red, 664.5nm (S2A) / 665nm (S2B)",
        "B8, NIR, 835.1nm (S2A) / 833nm (S2B)",
    ],
    time=["2026-01-01"],
)
image_again = rumi.read(image_path, image_header)
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

image_path identifies the durable file. image_header is the compact external index to keep with that object in a catalogue or dataset manifest.

Every file names each band and dates each time step. bands takes one text per band, in band order; the recommended text gives the band name, a short description and the wavelength.

Every frame is explicit

rumi.frames creates decoded frame views. The loop chooses an OpenZL graph and assigns the compressed payload for every frame. rumi.write refuses an incomplete table.

Choose what one frame contains

The logical Image is unchanged. Only the right side of the pattern changes how many samples must be fetched and decoded together.

# One independently addressable frame per band and spatial tile.
tile_frames = rumi.frames(
    image,
    "b (row h) (col w) -> row col b (h w)",
    tile_size=256,
)

# One band-planar cell frame at each spatial position.
planar_cells = rumi.frames(
    image,
    "b (row h) (col w) -> row col (b h w)",
    tile_size=256,
)

# One pixel-interleaved cell frame at each spatial position.
chunky_cells = rumi.frames(
    image,
    "b (row h) (col w) -> row col (h w b)",
    tile_size=256,
)
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

Frame layoutSelective-read behaviorCompression opportunity
h wUnrequested bands remain untouchedNo correlation across bands inside a frame
b h wSelecting one band decodes its complete cellBand planes share one OpenZL graph
h w bSelecting one band decodes its complete cellInterleaved pixels share one OpenZL graph

Write a temporal Cube

This Cube has three steps and four bands. Band and time sit outside each frame, so any individual pair can be read without decoding the others.

cube = np.random.default_rng(1).integers(
    0, 4096,
    size=(3, 4, 1024, 1024),
    dtype=np.uint16,
)

cube_frames = rumi.frames(
    cube,
    "t b (row h) (col w) -> row col b t (h w)",
    tile_size=256,
)

for frame in cube_frames:
    graph = geozl.graph(frame.data, "planar>zigzag>zstd")
    frame.compressed = geozl.compress(frame.data, graph=graph)

cube_path, cube_header = rumi.write(
    "cube.rumi",
    cube_frames,
    bands=[
        "B2, Blue, 496.6nm (S2A) / 492.1nm (S2B)",
        "B3, Green, 560nm (S2A) / 559nm (S2B)",
        "B4, Red, 664.5nm (S2A) / 665nm (S2B)",
        "B8, NIR, 835.1nm (S2A) / 833nm (S2B)",
    ],
    time=["2026-01-01", "2026-02-01", "2026-03-01"],
)
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

Tile-frame Cube

row col b t (h w) or row col t b (h w) keeps both axes independently addressable.

Cell-frame Cube

row col (b t h w) lets one graph model the entire band/time block at a spatial position.

The time argument accepts dates, datetimes, ISO strings, or one (start, end) interval per step. Its length must equal T.

Read positions, ranges, and windows

Lists choose exact positions and may reorder them. Two-item tuples describe half-open ranges. The optional output pattern changes tensor order without changing storage.

# Exact positions: time 2; bands 3 then 1.
picked = rumi.read(
    cube_path,
    cube_header,
    time=[2],
    bands=[3, 1],
    window=(100, 200, 128, 128),
)

# Half-open ranges and a model-specific output order.
reordered = rumi.read(
    cube_path,
    cube_header,
    time=(0, 2),
    bands=(1, 4),
    window=(100, 200, 128, 128),
    pattern="t y x b",
)
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

time=[2]

Keep exactly one step and preserve its time axis.

bands=[3, 1]

Keep two explicit bands in the requested order.

window=(100, 200, 128, 128)

Return 128 rows and columns beginning at row 100, column 200.

Read the same fixture from four sources

This example is ready to run against the public Rumi fixtures. It reads the same bands and window from Source Cooperative, Hugging Face, Source's S3-compatible endpoint, and a downloaded local file, then proves that all four results are identical.

from pathlib import Path
from urllib.request import Request, urlopen
import os

import numpy as np
import rumi

PRODUCT = "asterisk-labs/rumi-api-fixtures"
NAME = "s2-00-tile"
SOURCE_BASE = f"https://data.source.coop/{PRODUCT}"

def fetch(url):
    request = Request(url, headers={"User-Agent": "rumi-example"})
    with urlopen(request) as response:
        return response.read()

header = fetch(f"{SOURCE_BASE}/headers/{NAME}.header")
selection = dict(
    bands=[3, 0],
    window=(237, 233, 91, 107),
)

# Source Cooperative also exposes an S3-compatible public endpoint.
os.environ["AWS_ENDPOINT_URL"] = "https://data.source.coop"
os.environ["AWS_NO_SIGN_REQUEST"] = "YES"
os.environ["AWS_VIRTUAL_HOSTING"] = "NO"

sources = {
    "Source Cooperative": f"source://{PRODUCT}/data/{NAME}.rumi",
    "Hugging Face": f"hf://datasets/{PRODUCT}@main/data/{NAME}.rumi",
    "S3": f"s3://asterisk-labs/rumi-api-fixtures/data/{NAME}.rumi",
}
chips = {
    label: np.asarray(rumi.read(uri, header, **selection))
    for label, uri in sources.items()
}

# Download once to exercise the identical read from a local path.
local_dir = Path("rumi-api-fixtures")
local_dir.mkdir(exist_ok=True)
local_file = local_dir / f"{NAME}.rumi"
local_file.write_bytes(fetch(f"{SOURCE_BASE}/data/{NAME}.rumi"))
chips["Local"] = np.asarray(
    rumi.read(local_file, header, **selection)
)

reference = chips["Source Cooperative"]
assert reference.shape == (2, 91, 107)
assert all(np.array_equal(reference, chip) for chip in chips.values())

print({label: chip.shape for label, chip in chips.items()})
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

Source used aboveAddressConfiguration
Source Cooperativesource://asterisk-labs/rumi-api-fixtures/…None for public data
Hugging Facehf://datasets/asterisk-labs/rumi-api-fixtures@main/…None for this public dataset
S3-compatibles3://asterisk-labs/rumi-api-fixtures/…Source endpoint + unsigned requests
Localrumi-api-fixtures/s2-00-tile.rumiDownloaded by the example
One object, four transports

The two mirrors contain byte-identical Rumi objects and headers. The S3 URI addresses the same Source Cooperative object as the source:// URI; only the backend configuration changes. Remote reads request ranges, while the local read uses positional file access.

Read one training window per source

read_many plans the items together and returns them under a leading n axis. Positions may differ, but every window must have the same height and width.

# Continue from the previous example.
names = ["s2-00-tile", "s2-01-tile", "s2-02-tile"]
headers = [
    fetch(f"{SOURCE_BASE}/headers/{name}.header")
    for name in names
]
sources = [
    f"source://{PRODUCT}/data/{names[0]}.rumi",
    f"hf://datasets/{PRODUCT}@main/data/{names[1]}.rumi",
    f"s3://asterisk-labs/rumi-api-fixtures/data/{names[2]}.rumi",
]
windows = [
    (0, 0, 128, 128),
    (128, 192, 128, 128),
    (320, 256, 128, 128),
]

batch = rumi.read_many(
    sources,
    headers,
    windows=windows,
    bands=[0, 1, 2, 3],
)

assert batch.shape == (3, 4, 128, 128)
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

Sources may mix local and remote storage. They must agree on tile size, dtype, band/time counts, and which axes their frames contain. The result preserves input order even when ranges complete in another order.

Why not a Python loop?

The batch exposes independent work from every item at once, allowing remote waits, local positional reads, and OpenZL decoding to overlap while Rumi assembles one predictable tensor.

Return the tensor your pipeline expects

Rumi decodes on the CPU. Select the destination at the read boundary so most results transfer through DLPack without an intermediate NumPy copy.

numpy_chip = rumi.read(cube_path, cube_header)
torch_chip = rumi.read(
    cube_path, cube_header, framework="torch"
)
jax_chip = rumi.read(
    cube_path, cube_header, framework="jax"
)
tensorflow_chip = rumi.read(
    cube_path, cube_header, framework="tensorflow"
)
dlpack_chip = rumi.read(
    cube_path, cube_header, framework="dlpack"
)
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

frameworkReturned value
"numpy"NumPy array; the default
"torch"PyTorch tensor
"jax"JAX array; 64-bit dtypes require jax_enable_x64
"tensorflow"TensorFlow tensor
"dlpack"RumiArray, a one-shot DLPack producer

Every Rumi dtype maps exactly to a CPU dtype accepted by PyTorch. NumPy, JAX and TensorFlow accept exact subsets. Rumi checks compatibility before opening the source and never casts or returns a byte view.

Inspect a file and recover its header

A Rumi file remains self-describing. If the external header is missing, info reconstructs the canonical header and also exposes the band texts, time and georeferencing stored in the file.

metadata = rumi.info(source="cube.rumi")
recovered_header = metadata.header

print(metadata.shape)         # (3, 4, 1024, 1024)
print(metadata.tile)          # (256, 256)
print(metadata.frame_layout)  # h w for the tile-frame Cube
print(metadata.bands)         # the four band texts
print(metadata.time)          # the three recorded dates

# Supplying both validates the index against the source.
checked = rumi.info(
    source="cube.rumi",
    header=recovered_header,
)
RComing soon

Native R bindings are not available yet.

JuliaComing soon

Native Julia bindings are not available yet.

Header only

rumi.info(header=...) describes the compact read index without opening the source.

Source included

rumi.info(source=...) also returns band texts, time coordinates, transform, CRS, and a rebuilt header.

Understand the execution path in How Rumi Reads, open the format specification, or return to the Rumi overview.