Reference

Public API

class timedb.TimeDBClient(ch_url: str | None = None)[source]

Bases: object

__init__(ch_url: str | None = None)[source]
close() → None[source]

Close the underlying ClickHouse connection.

create() → None[source]

Create the series_values table and run_series mapping.

delete() → None[source]

Drop both CH tables.

read(*, series_ids: Sequence[int], retention: str | Sequence[str] | None = None, start_valid: datetime | None = None, end_valid: datetime | None = None, start_known: datetime | None = None, end_known: datetime | None = None, include_updates: bool = False, include_knowledge_time: bool = False, meta_source: PgEngineMeta | None = None) → DataFrame[source]

Read values for series_ids, returning a Polars DataFrame.

By default this collapses to the latest value per valid_time, the row with the largest (knowledge_time, change_time), and returns series_id, valid_time, value. Two flags widen it:

  • include_knowledge_time=True: one row per (knowledge_time, valid_time), every forecast run side by side, adding knowledge_time.

  • include_updates=True: the full correction chain on the winning run, adding change_time, changed_by and annotation.

Setting both returns the complete 3-dimensional audit log.

retention accepts one tier or a sequence of tiers and prunes whole partitions. start_valid / end_valid bound valid_time; start_known / end_known bound knowledge_time. All datetimes must be timezone-aware.

meta_source takes a PgEngineMeta to have ClickHouse resolve the series set itself through a PostgreSQL engine table instead of receiving an explicit id array. energydb’s concurrent read path uses it.

read_relative(*, series_ids: Sequence[int], retention: str | Sequence[str] | None = None, window_length: timedelta | None = None, issue_offset: timedelta | None = None, start_window: datetime | None = None, start_valid: datetime | None = None, end_valid: datetime | None = None, days_ahead: int | None = None, time_of_day: time | None = None, meta_source: PgEngineMeta | None = None) → DataFrame[source]

Per-window cutoff read: for each window, the latest forecast issued at or before that window’s cutoff.

This is the “what forecast was available at decision time” read that backtests and day-ahead simulations need. Returns series_id, valid_time, value.

Two mutually exclusive parameter sets address the windows; mixing them raises ValueError:

  • Low-level: window_length plus issue_offset (relative to each window start) and start_window.

  • Daily shorthand: days_ahead plus time_of_day, giving fixed 1-day windows with a human-friendly cutoff.

start_valid / end_valid bound the returned range, retention prunes partitions, and meta_source behaves as in read(). All datetimes must be timezone-aware.

read_run_series(*, series_id: int) → list[int][source]

Return run_ids that touched a given series_id, latest first.

Data only: the energydb.runs PG table hydrates the metadata.

write(df: DataFrame | DataFrame, *, retention: str | None = None, knowledge_time: datetime | None = None, skip_unchanged: bool = False, unchanged_scope: Literal['valid_time', 'knowledge_time', 'auto'] = 'valid_time', knowledge_time_scoped_series: Collection[int] | None = None) → WriteResult[source]

Write time-series rows into series_values and their run_series mapping.

df may be a Pandas or Polars frame. Required columns: series_id, valid_time, value. Optional columns get a per-batch default when absent: knowledge_time (this kwarg, else datetime.now(UTC)), change_time (now(UTC)), run_id (one client-generated UUID7 truncated to 63 bits), valid_time_end (the 2200-01-01 sentinel), and changed_by / annotation (empty strings). Every timestamp column must be timezone-aware; naive values raise ValueError.

retention and knowledge_time may be given as a kwarg or a column, never both. retention defaults to "forever" (no TTL); see RETENTION_TIERS for the valid tiers.

With skip_unchanged=True, rows whose latest stored (value, annotation, changed_by) already matches are dropped before the insert, at the cost of one bounded read-back. unchanged_scope picks the comparison key: "valid_time" (default), "knowledge_time", or "auto", which applies the knowledge-time key to the ids in knowledge_time_scoped_series and the valid-time key to every other series. Any other scope paired with knowledge_time_scoped_series raises.

Returns a WriteResult: a NamedTuple(written, skipped) of row counts.

timedb.RETENTION_TIERS = frozenset({'forever', 'long', 'medium', 'short'})

frozenset() -> empty frozenset object frozenset(iterable) -> frozenset object

Build an immutable unordered collection of unique elements.

class timedb.WriteResult(written: int, skipped: int)[source]

Bases: NamedTuple

Counts returned by write(). skipped is always 0 unless skip_unchanged was set.

skipped: int

Alias for field number 1

written: int

Alias for field number 0

timedb.UnchangedScope

The comparison key for write(skip_unchanged=True): "valid_time", "knowledge_time", or "auto" (per-series, driven by knowledge_time_scoped_series). alias of Literal[‘valid_time’, ‘knowledge_time’, ‘auto’]

class timedb.PgEngineMeta(table: str, root_path: str | None = None, paths: tuple[str, ...] | None = None, node_uuids: tuple[str, ...] | None = None, edge_uuids: tuple[str, ...] | None = None, edge_triple: tuple[str, str, str] | None = None, edge_triples: tuple[tuple[str, str, str], ...] | None = None, data_type: str | tuple[str, ...] | None = None, name: str | tuple[str, ...] | None = None)[source]

Bases: object

Resolve the series_id set inside ClickHouse via a PostgreSQL engine table over the series_meta view, instead of the caller passing an explicit id array. The read then filters series_id and retention by subqueries over this meta CTE rather than array parameters.

Exactly one addressing field must be set:

  • root_path: a node subtree (the root itself + descendants, path-prefix match);

  • paths: an exact set of node paths (path-routed manifests);

  • node_uuids / edge_uuids: owner-uuid sets (uuid-routed manifests, edge scopes);

  • edge_triple: one (from_path, to_path, edge_type) edge identity (edge scopes);

  • edge_triples: a set of (from_path, to_path, edge_type) identities (triple-routed manifests), pushed down as three single-column IN filters. Like the set-valued data_type / name below, this resolves a cartesian superset of the requested triples; the caller trims against its exactly-resolved meta.

data_type / name narrow the series set; each accepts a scalar (scope reads) or a set of values (manifests). Set-valued filters make the engine-resolved ids a superset (the cartesian of the sets). Every predicate here pushes down to PG as a single-column comparison, and the caller is expected to trim against its exactly-resolved meta.

Profiling helpers

A lightweight phase-timer used by the read/write paths. Useful when diagnosing slow queries or large bulk inserts.

Opt-in per-phase timing collector for TimeDB internal operations.

Disabled by default: zero overhead when disabled (no perf_counter calls, no function calls in the hot path). Benchmark scripts activate it per-trial to collect phase-level timing breakdowns.

Not thread-safe; designed for single-threaded benchmark use.

Usage:

from timedb import profiling

profiling.enable()
profiling.reset()
# ... run operation ...
phases = profiling.collect()   # dict of phase -> elapsed seconds
profiling.disable()

Or, for hot-path instrumentation:

with profiling._phase(profiling.PHASE_EDB_RESOLVE):
    ...
timedb.profiling.collect() → dict[str, float][source]

Return a copy of accumulated timings (in seconds).

timedb.profiling.disable() → None[source]

Disable profiling and clear timings.

timedb.profiling.enable() → None[source]

Enable profiling collection. Call before each trial.

timedb.profiling.is_enabled() → bool[source]

Return True if profiling is currently active.

timedb.profiling.reset() → None[source]

Clear accumulated timings while keeping profiling enabled.