ssb_timeseries.io.archiving

Archiving: where an archive goes, and what is kept of it.

This module holds only the parts of archiving that are archiving rather than convention. Where the file goes, what it is called and how its bytes are written are all conventions, and belong to ssb_timeseries.io.format. See ssb_timeseries.io.archive() for archiving a dataset.

class Archive(repository, path='', archive_format=None, **options)

Bases: object

Writes a versioned copy of a dataset’s data, and keeps every version.

An archive is a destination, so it is told what to archive rather than asked for a path. A dataset is archived wherever its data repository keeps it, which need not be a filesystem at all.

Parameters:
  • repository (str | dict)

  • path (PathStr)

  • archive_format (str | ArchiveFormat | None)

  • options (Any)

__init__(repository, path='', archive_format=None, **options)

Initialize the archive handler from the archive’s configuration.

Which dataset to archive is passed to each operation as a DatasetRef, so one instance can archive any number of datasets.

Parameters:
  • repository (str | dict) – The archive repository name or configuration.

  • path (str | PathLike[str]) – The root of the archive.

  • archive_format (str | ArchiveFormat | None) – The name of the convention to follow, or a convention itself. Defaults to the convention in GENERIC.

  • **options (typing.Any) – Any further parameters defined for the handler in the configuration.

Return type:

None

directory(ref, tokens)

Return the folder a dataset’s archives belong in, creating it.

Parameters:
  • ref (DatasetRef) – The dataset being archived.

  • tokens (dict[str, str]) – The naming tokens, which carry the folder values.

Return type:

str | PathLike[str]

Returns:

The folder path.

naming_tokens(ref, period_from=None, period_to=None)

Collect the tokens the convention names a file by.

Parameters:
  • ref (DatasetRef) – The dataset being archived.

  • period_from (datetime | date | None) – The start of the data’s time period, if it has one.

  • period_to (datetime | date | None) – The end of the data’s time period, if it has one.

Return type:

dict[str, str]

Returns:

The tokens, with empty ones rendered as empty strings so that a template may interpolate them unconditionally.

next_version_number(ref, directory)

Find the version number to give the next archive of a dataset.

Every version already archived is counted, because the convention is that an archive is never overwritten and never removed.

Parameters:
  • ref (DatasetRef) – The dataset being archived.

  • directory (str | PathLike[str]) – The folder its archives are in.

Return type:

int

Returns:

The version number to use.

replicate(ref, file_path, destinations, relative_folder='')

Copy an archived file to further destinations.

The archive is written once and copied from, rather than written again per destination, so that replication costs a copy rather than a re-encoding of the data.

Each destination gets the archive at the same place under it as it has under the archive’s own root, so that a location holding several datasets still resolves them the same way the archive does.

Parameters:
  • ref (DatasetRef) – The dataset that was archived, for logging.

  • file_path (str | PathLike[str]) – The file to copy.

  • destinations (list[str | PathLike[str]]) – The folders to copy it under.

  • relative_folder (str | PathLike[str]) – The dataset’s folder, relative to the archive root.

Return type:

None

write(ref, data, period_from=None, period_to=None, destinations=None)

Write one version of a dataset’s data, and keep every earlier one.

The data is read by the caller rather than copied from a path, so that a dataset which is not stored as a file can be archived just as well. No version is ever overwritten: each call adds a new file.

Parameters:
  • ref (DatasetRef) – The dataset being archived.

  • data (typing.Any) – The dataset’s data, as read from the data repository.

  • period_from (datetime | date | None) – The start of the data’s time period, if it has one.

  • period_to (datetime | date | None) – The end of the data’s time period, if it has one.

  • destinations (list[str | PathLike[str]] | None) – Further folders to copy the archive into. Only used when the convention requires shared data to be archived.

Return type:

str | PathLike[str]

Returns:

The path of the archive that was written.