Sitelet https://github.com/kitsuyui/python-tally-token
Skip to content

python-tally-token

Python PyPI version shields.io License Coverage

What is this?

tally-token is a Python library for split data into tokens with same length.

Tally is a historical object for prove something by splitting wood into tokens and matching tokens.

Medieval English split tally stick (front and reverse view). The stick is notched and inscribed to record a debt owed to the rural dean of Preston Candover, Hampshire, of a tithe of 20d each on 32 sheep, amounting to a total sum of £2 13s. 4d.

Changelog

See CHANGELOG.md for release notes and upgrade notes.

Usage

Install

$ pip install tally-token

CLI Usage

$ tally-token --help
usage: tally-token [-h] {split,merge} ...

positional arguments:
  {split,merge}  Commands: split: split a file into multiple files merge: merge multiple files into a fileExample: tally-token split example.bin
                 example.bin.1 example.bin.2 example.bin.3 tally-token merge example.bin.1 example.bin.2 example.bin.3 --output example-merged.bin

options:
  -h, --help     show this help message and exit

split

You can use split to split a file into multiple files.

$ tally-token split something.bin split-1.bin split-2.bin split-3.bin

merge

You can use merge to merge multiple files into a file.

$ tally-token merge split-1.bin split-2.bin split-3.bin --output merged.bin

Large files

Nothing special. You can split and merge large file.

$ dd if=/dev/urandom of=original.1g.bin bs=1G count=1
$ tally-token split original.1g.bin split-1.bin split-2.bin split-3.bin
$ shasum -a 256 original.1g.bin
> 736a344d99d27e2dcdab8bc37ca94c83eda26f812a3dee87ac98989f89b3f965 original.1g.bin
$ tally-token merge split-1.bin split-2.bin split-3.bin --output recovery.1g.bin
$ shasum -a 256 recovery.1g.bin
> 736a344d99d27e2dcdab8bc37ca94c83eda26f812a3dee87ac98989f89b3f965 recovery.1g.bin

Example

split

You can use split_text to split text into tokens. split_text returns list of random bytes.

>>> from tally_token import split_text
>>> split_text("Hello, World!")
[b'qQ\xa5\x97\x84\x88\xd7U%\xfb(k\xa1', b'94\xc9\xfb\xeb\xa4\xf7\x02J\x89D\x0f\x80']

merge

You can use merge_text to merge tokens into text. merge_text returns cleartext.

>>> from tally_token import merge_text
>>> merge_text([b'qQ\xa5\x97\x84\x88\xd7U%\xfb(k\xa1', b'94\xc9\xfb\xeb\xa4\xf7\x02J\x89D\x0f\x80'])
'Hello, World!'

split with custom length

>>> from tally_token import split_text, merge_text
>>> split_text("Hello, World!", 5)
[b'N&\xce\\\xbc6dxp\x87\xa8#z', b'\xa3D\\A\xf8\xd1KDX\x1cKx\x87', b'\xffZ\x03\xf5\x92Q\xf52\xc4\x1e\xf2\xf8\x06', b'\xaa\xdd:\x85F\xa1\xcdbp\xf3\xe6P\xe5', b'\xf0\x80\xc7\x01\xff;7;\xf3\x04\x9b\x97?']
>>> merge_text([b'N&\xce\\\xbc6dxp\x87\xa8#z', b'\xa3D\\A\xf8\xd1KDX\x1cKx\x87', b'\xffZ\x03\xf5\x92Q\xf52\xc4\x1e\xf2\xf8\x06', b'\xaa\xdd:\x85F\xa1\xcdbp\xf3\xe6P\xe5', b'\xf0\x80\xc7\x01\xff;7;\xf3\x04\x9b\x97?'])
'Hello, World!'

split with custom encoding

>>> from tally_token import split_text, merge_text
>>> split_text("こんにちは", encoding="CP932")
[b'g\xc3\x12\xeal?\xe5[\x03\xad', b'\xe5r\x90\x1b\xee\xf6g\xe4\x81`']
>>> merge_text([b'g\xc3\x12\xeal?\xe5[\x03\xad', b'\xe5r\x90\x1b\xee\xf6g\xe4\x81`'], encoding="CP932")
'こんにちは'

The returned tokens are opaque bytes and do not store the encoding name. When splitting text with a custom encoding, keep that encoding out of band and pass the same value to merge_text; using a different encoding can raise UnicodeDecodeError or return mojibake if the bytes happen to decode.

bytes interface

You can use split_bytes_into and merge_bytes_into to split and merge bytes. This is useful for split binary data.

>>> from tally_token import split_bytes_into, merge_bytes_into
>>> split_bytes_into(b"Hello, World!", 5)
[b'\xc5b\xf4E)\xe1vO8\xff@\xf9\xdd', b'\x84\xb9X#\x85\xf5\xed\xbcM\xc4\xef\xf4\xd3', b'\xb47\xf6\xfa?\x14\xa8`\xc9\xe0\xe5\x87\x14', b'\x1cd\xb4o\xe8I:\xe5\xf6\x13\xe5\x93G', b'\xa1\xed\x82\x9f\x14e)!%\xba\xc3}|']
>>> merge_bytes_into([b'\xc5b\xf4E)\xe1vO8\xff@\xf9\xdd', b'\x84\xb9X#\x85\xf5\xed\xbcM\xc4\xef\xf4\xd3', b'\xb47\xf6\xfa?\x14\xa8`\xc9\xe0\xe5\x87\x14', b'\x1cd\xb4o\xe8I:\xe5\xf6\x13\xe5\x93G', b'\xa1\xed\x82\x9f\x14e)!%\xba\xc3}|'])
b'Hello, World!'

Security and Logging

Security model

This library implements an n-of-n XOR split. All tokens produced by the same split operation are required to recover the original bytes. It is not a k-of-n threshold scheme such as Shamir's Secret Sharing, so any missing token prevents recovery.

Tokens include a shared session ID so merge_bytes_into and merge_io can reject tokens from different split operations. That check is not a message authentication or tamper-detection mechanism: if a token from the correct session is modified, merging can still return modified bytes without raising an error.

Calling split_bytes_into(secret, 1) or split_text(secret, into=1) does not protect confidentiality. The single returned token contains the original secret payload after the session ID, so use at least two tokens when the goal is to split secret material.

Report vulnerabilities through the private channel described in SECURITY.md.

Logging policy

This library handles secret material (token bytes) and follows a log-nothing policy: no secret data should appear in log output.

  • The library installs a NullHandler on its root logger so log records are silently discarded unless the application configures a handler. This is the standard practice for libraries.
  • Do not log token bytes or plain-text secrets in your application code. Functions such as split_bytes_into, _split1, write_split_tokens, and merge_bytes_into operate directly on secret material.
  • If you add debug logging to code that calls this library, ensure no token content is captured in log records or exception messages.

Reference

Development

This repository uses lefthook to run the same checks as CI locally, so problems surface before they reach CI.

# Install dependencies
uv sync

# Install the Git hooks (once; requires lefthook on your PATH)
lefthook install

Once installed, the hooks run automatically:

  • pre-commit: uv run poe check
  • pre-push: uv run poe check and uv run poe test

You can also run the checks manually:

uv run poe check
uv run poe test

CI still runs the full matrix (see .github/workflows/); the hooks only bring that feedback earlier on your machine.

LICENSE

BSD 3-Clause License

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages