tally-token is a Python library for split data into tokens with same length.
Tally is a historical object for prove something by splitting wood into tokens and matching tokens.
Medieval English split tally stick (front and reverse view). The stick is notched and inscribed to record a debt owed to the rural dean of Preston Candover, Hampshire, of a tithe of 20d each on 32 sheep, amounting to a total sum of £2 13s. 4d.
See CHANGELOG.md for release notes and upgrade notes.
$ pip install tally-token$ tally-token --help
usage: tally-token [-h] {split,merge} ...
positional arguments:
{split,merge} Commands: split: split a file into multiple files merge: merge multiple files into a fileExample: tally-token split example.bin
example.bin.1 example.bin.2 example.bin.3 tally-token merge example.bin.1 example.bin.2 example.bin.3 --output example-merged.bin
options:
-h, --help show this help message and exitYou can use split to split a file into multiple files.
$ tally-token split something.bin split-1.bin split-2.bin split-3.binYou can use merge to merge multiple files into a file.
$ tally-token merge split-1.bin split-2.bin split-3.bin --output merged.binNothing special. You can split and merge large file.
$ dd if=/dev/urandom of=original.1g.bin bs=1G count=1
$ tally-token split original.1g.bin split-1.bin split-2.bin split-3.bin
$ shasum -a 256 original.1g.bin
> 736a344d99d27e2dcdab8bc37ca94c83eda26f812a3dee87ac98989f89b3f965 original.1g.bin
$ tally-token merge split-1.bin split-2.bin split-3.bin --output recovery.1g.bin
$ shasum -a 256 recovery.1g.bin
> 736a344d99d27e2dcdab8bc37ca94c83eda26f812a3dee87ac98989f89b3f965 recovery.1g.binYou can use split_text to split text into tokens. split_text returns list of random bytes.
>>> from tally_token import split_text
>>> split_text("Hello, World!")
[b'qQ\xa5\x97\x84\x88\xd7U%\xfb(k\xa1', b'94\xc9\xfb\xeb\xa4\xf7\x02J\x89D\x0f\x80']You can use merge_text to merge tokens into text. merge_text returns cleartext.
>>> from tally_token import merge_text
>>> merge_text([b'qQ\xa5\x97\x84\x88\xd7U%\xfb(k\xa1', b'94\xc9\xfb\xeb\xa4\xf7\x02J\x89D\x0f\x80'])
'Hello, World!'>>> from tally_token import split_text, merge_text
>>> split_text("Hello, World!", 5)
[b'N&\xce\\\xbc6dxp\x87\xa8#z', b'\xa3D\\A\xf8\xd1KDX\x1cKx\x87', b'\xffZ\x03\xf5\x92Q\xf52\xc4\x1e\xf2\xf8\x06', b'\xaa\xdd:\x85F\xa1\xcdbp\xf3\xe6P\xe5', b'\xf0\x80\xc7\x01\xff;7;\xf3\x04\x9b\x97?']
>>> merge_text([b'N&\xce\\\xbc6dxp\x87\xa8#z', b'\xa3D\\A\xf8\xd1KDX\x1cKx\x87', b'\xffZ\x03\xf5\x92Q\xf52\xc4\x1e\xf2\xf8\x06', b'\xaa\xdd:\x85F\xa1\xcdbp\xf3\xe6P\xe5', b'\xf0\x80\xc7\x01\xff;7;\xf3\x04\x9b\x97?'])
'Hello, World!'>>> from tally_token import split_text, merge_text
>>> split_text("こんにちは", encoding="CP932")
[b'g\xc3\x12\xeal?\xe5[\x03\xad', b'\xe5r\x90\x1b\xee\xf6g\xe4\x81`']
>>> merge_text([b'g\xc3\x12\xeal?\xe5[\x03\xad', b'\xe5r\x90\x1b\xee\xf6g\xe4\x81`'], encoding="CP932")
'こんにちは'The returned tokens are opaque bytes and do not store the encoding name. When
splitting text with a custom encoding, keep that encoding out of band and pass
the same value to merge_text; using a different encoding can raise
UnicodeDecodeError or return mojibake if the bytes happen to decode.
You can use split_bytes_into and merge_bytes_into to split and merge bytes.
This is useful for split binary data.
>>> from tally_token import split_bytes_into, merge_bytes_into
>>> split_bytes_into(b"Hello, World!", 5)
[b'\xc5b\xf4E)\xe1vO8\xff@\xf9\xdd', b'\x84\xb9X#\x85\xf5\xed\xbcM\xc4\xef\xf4\xd3', b'\xb47\xf6\xfa?\x14\xa8`\xc9\xe0\xe5\x87\x14', b'\x1cd\xb4o\xe8I:\xe5\xf6\x13\xe5\x93G', b'\xa1\xed\x82\x9f\x14e)!%\xba\xc3}|']
>>> merge_bytes_into([b'\xc5b\xf4E)\xe1vO8\xff@\xf9\xdd', b'\x84\xb9X#\x85\xf5\xed\xbcM\xc4\xef\xf4\xd3', b'\xb47\xf6\xfa?\x14\xa8`\xc9\xe0\xe5\x87\x14', b'\x1cd\xb4o\xe8I:\xe5\xf6\x13\xe5\x93G', b'\xa1\xed\x82\x9f\x14e)!%\xba\xc3}|'])
b'Hello, World!'This library implements an n-of-n XOR split. All tokens produced by the same split operation are required to recover the original bytes. It is not a k-of-n threshold scheme such as Shamir's Secret Sharing, so any missing token prevents recovery.
Tokens include a shared session ID so merge_bytes_into and merge_io can
reject tokens from different split operations. That check is not a message
authentication or tamper-detection mechanism: if a token from the correct
session is modified, merging can still return modified bytes without raising an
error.
Calling split_bytes_into(secret, 1) or split_text(secret, into=1) does not
protect confidentiality. The single returned token contains the original secret
payload after the session ID, so use at least two tokens when the goal is to
split secret material.
Report vulnerabilities through the private channel described in SECURITY.md.
This library handles secret material (token bytes) and follows a log-nothing policy: no secret data should appear in log output.
- The library installs a
NullHandleron its root logger so log records are silently discarded unless the application configures a handler. This is the standard practice for libraries. - Do not log token bytes or plain-text secrets in your application code.
Functions such as
split_bytes_into,_split1,write_split_tokens, andmerge_bytes_intooperate directly on secret material. - If you add debug logging to code that calls this library, ensure no token content is captured in log records or exception messages.
- Tally stick https://en.wikipedia.org/wiki/Tally_stick
- Tessera or Symbolum (Hospitium token) https://en.wikipedia.org/wiki/Hospitium
- 割符 https://ja.wikipedia.org/wiki/%E5%89%B2%E7%AC%A6
- One-time pad https://en.wikipedia.org/wiki/One-time_pad
This repository uses lefthook to run the same checks as CI locally, so problems surface before they reach CI.
# Install dependencies
uv sync
# Install the Git hooks (once; requires lefthook on your PATH)
lefthook installOnce installed, the hooks run automatically:
- pre-commit:
uv run poe check - pre-push:
uv run poe checkanduv run poe test
You can also run the checks manually:
uv run poe check
uv run poe testCI still runs the full matrix (see .github/workflows/); the hooks only bring that
feedback earlier on your machine.
BSD 3-Clause License
