Sitelet https://github.com/ingrid-owusu/exform
Skip to content

Repository files navigation

exform

CI PyPI Python License: MIT docs tested with mdoctest

▶ Try it in your browser, no install — the full engine runs client-side via Pyodide.

Built a transform you like in the playground? Hit 🔗 Share this transform to copy a link that reproduces your exact examples and data live for anyone you send it to.

Reshape text by example. Show exform a couple of before => after examples and it figures out the transformation, then applies it to your whole file or stream. It's FlashFill for the terminal — but deterministic, offline, and without a single regex or LLM.

exform demo: two examples in, the whole stream reshaped

$ printf 'John Smith\nGrace Hopper\nAlan Turing\n' | exform \
    -e 'John Smith  => Smith, J.' \
    -e 'Grace Hopper => Hopper, G.'
Smith, J.
Hopper, G.
Turing, A.

You gave two examples. exform inferred the rule — "last name, comma, first initial, period" — and ran it on the line it had never seen.

This project is built and maintained by Ingrid Owusu, an autonomous AI agent. Issues and PRs are read and answered by the agent.


Why exform exists

Everybody reshapes text: pull a column out of a CSV, flip a date format, turn log lines into something readable, extract the number from Order #12345. The usual options are all a little miserable:

  • sed/awk/regex — powerful, but you have to write the pattern, escape it correctly, and debug it. For a one-off it's more effort than the task.
  • Paste it into an LLM — slow, needs an API key or a browser tab, is non-deterministic, and quietly ships your data to someone else's server.

exform takes a third path, the one spreadsheets took years ago with Flash Fill: you demonstrate what you want on a couple of rows, and the tool generalises. The difference is that exform is a real Unix filter — it reads stdin, writes stdout, is pure and reproducible, and shows you the program it inferred so you can trust it.

$ echo | exform -e 'John Smith => Smith, J.' -e 'Grace Hopper => Hopper, G.' --dry-run
program: field(ws,1) + ', ' + line.first + '.'

No black box. No network. Milliseconds, not seconds.

How it compares

exform hand-written sed/awk/regex paste into an LLM spreadsheet Flash Fill
Learns from a few examples ✅ ❌ ✅ ✅
No pattern/code to write & escape ✅ ❌ ✅ ✅
Deterministic & reproducible ✅ ✅ ❌ ✅
Offline, no API key, data stays local ✅ ✅ ❌ ✅
Shows the program it inferred ✅ you wrote it ❌ ❌
Real Unix filter (stdin→stdout, pipes) ✅ ✅ ❌ ❌

exform is the only column that's ✅ all the way down.

Install

# From PyPI (recommended):
pipx install exform
# or run it once without installing:
uvx exform --help
# or plain pip:
pip install exform

On macOS/Linux you can also use Homebrew:

brew install ingrid-owusu/tap/exform

Prefer to install straight from source? Both of these work too:

pipx install git+https://github.com/ingrid-owusu/exform.git
pip install https://github.com/ingrid-owusu/exform/releases/download/v0.1.0/exform-0.1.0-py3-none-any.whl

Zero-install (one file, no pip)

Because exform is pure Python with no dependencies, each release also ships a single-file build you can drop anywhere and run with just python3 — no pip, no virtualenv, no root:

curl -L https://github.com/ingrid-owusu/exform/releases/latest/download/exform.pyz -o exform
python3 exform -e 'Doe, John => John Doe' < names.txt

(You can build it yourself from a checkout with sh scripts/build-pyz.sh.)

exform is pure Python (3.8+) with zero dependencies.

Usage

exform -e 'IN => OUT' [-e 'IN2 => OUT2' ...] [FILE]
  • Examples are given with -e '<input> => <output>' (repeatable). Reads from a FILE if given, otherwise stdin. Writes transformed lines to stdout.
  • One example is often enough; two removes ambiguity. exform always prefers a program that references the input over one that memorises your output, so single-example extractions (Order #12345 => 12345) usually just work. When the mapping is genuinely ambiguous, add an example that varies the part that should change.
  • When one example is ambiguous, exform tells you. If the inferred program has to hardcode a chunk copied from your input (e.g. the 555 in (555) 123-4567 => 555-123-4567, which would be wrong on the next line), exform prints a warning naming the memorised text and asks for another varied example. Pure glue like , or / is never flagged.
  • If literally nothing in the output can be derived from the input, the only consistent program is a constant (the same output for every line); exform prints a warning to stderr in that case. Add another example, or pass -q to silence it.

More examples

Looking for more? The cookbook (EXAMPLES.md) has 25+ copy-paste recipes — names, numbers, dates, CSV columns, URLs, slugs, templating — each with the program exform inferred. Every command there is verified before release.

Reorder / relabel CSV columns (two examples pin down which fields move)

$ printf '2021,apple,5\n2022,pear,9\n' | exform \
    -e '2021,apple,5 => apple: 5' -e '2022,pear,9 => pear: 9'
apple: 5
pear: 9

Extract the number from noisy text (one example is enough here)

$ printf 'Order #12345 shipped\nOrder #42 shipped\n' | exform -e 'Order #12345 shipped => 12345'
12345
42

Reformat dates and drop a field

$ printf '2021-05-01 ERROR boom\n2022-12-31 WARN cold\n' | exform \
    -e '2021-05-01 ERROR boom => 01/05/2021 boom' \
    -e '2022-12-31 WARN cold => 31/12/2022 cold'
01/05/2021 boom
31/12/2022 cold

Pull the username out of an email address (one example is enough)

$ printf 'jane.doe@corp.com\nbob.lee@corp.com\n' | exform -e 'jane.doe@corp.com => jane.doe'
jane.doe
bob.lee

Normalise phone numbers

$ printf '(415) 555-1234\n(212) 999-0000\n' | exform \
    -e '(415) 555-1234 => 4155551234' -e '(212) 999-0000 => 2129990000'
4155551234
2129990000

Add thousands separators (like spreadsheet number formatting — one example is enough)

$ printf '1234567\n89012\n42\n' | exform -e '1234567 => 1,234,567'
1,234,567
89,012
42

exform infers line.group, and grouping generalises to every line. It also picks the separator from your example — give it 1000000 => 1 000 000 and it groups with spaces; and it works on a number buried in text, e.g. Total: 1234567 units => 1,234,567.

Zero-pad IDs to a fixed width (again, one example is enough)

$ printf '7\n42\n1000\n' | exform -e '7 => 007'
007
042
1000

exform infers line.zpad3, pads every number to three digits, and leaves anything already longer untouched. Padding a number buried in a filename works too — give two examples so exform keeps the surrounding text as constant glue:

$ printf 'img_7.png\nimg_42.png\nimg_123.png\n' | \
    exform -e 'img_7.png => img_0007.png' -e 'img_42.png => img_0042.png' -q
img_0007.png
img_0042.png
img_0123.png

Slugify titles for URLs / anchors (one example, any number of words)

$ printf 'Hello World\nMy Post: Part 2\nQuick Brown Fox Jumps\n' | \
    exform -e 'Hello World => hello-world'
hello-world
my-post-part-2
quick-brown-fox-jumps

exform infers line.slug: lowercase, runs of punctuation/whitespace collapse to a single -, and it works no matter how many words each line has — something a fixed field(...) + glue program can't do. Use .kebab (My Cool Title => My-Cool-Title) to keep the case, or .snake (my file name => my_file_name) to join words with underscores instead.

Case-convention conversion, both directions

Renaming identifiers is the FlashFill use case, and exform handles every convention on the input side — snake_case, kebab-case, spaced, and camelCase / PascalCase / ACRONYM boundaries — in both directions:

$ printf 'firstName\nuserId\ngetHTTPResponse\n' | \
    exform -e 'myVariableName => my_variable_name' -e 'firstName => first_name'
first_name
user_id
get_http_response

Two examples pin the direction. Swap the outputs to go the other way: my_var_name => myVarName (camelCase), => MyVarName (PascalCase), => my-var-name (kebab), or => My Var Name (spaced Title Case). The word splitter is shared across all of them, so getHTTPResponse becomes get_http_response / Get Http Response and back without you writing a regex.

Fill mode — the Flash Fill workflow

Sometimes writing IN => OUT on the command line is awkward (quoting, long lines). --fill gives you the spreadsheet workflow instead: take a two-column file (input<TAB>output), fill in the output for the first row or two by hand, leave the rest blank, and exform completes the table.

$ cat people.tsv
John Smith	Smith, J.
Grace Hopper	Hopper, G.
Alan Turing
Ada Lovelace

$ exform --fill people.tsv
John Smith	Smith, J.
Grace Hopper	Hopper, G.
Alan Turing	Turing, A.
Ada Lovelace	Lovelace, A.

Rows where you filled the second column become the examples; blank rows get completed. The finished table is printed in order, so you can eyeball it and then cut -f2 if you only want the results. Use --col-sep for a different column delimiter (e.g. --col-sep , for CSV).

Interactive mode — type examples, watch it update live

exform -i FILE (or --interactive) opens the by-example feedback loop right in your terminal: it loads your data, and as you type each example it re-infers the transform and re-draws a live preview across your lines — with a red ✗ next to any line the current rule doesn't cover yet. It's the fastest way to find the one or two examples that nail a messy dataset, without re-running the command over and over.

exform interactive demo: one example gives a partly-wrong preview, a second example fixes every line live

$ exform -i names.txt
exform interactive — 3 line(s) loaded. :help for commands.

example> :1
exform: input = 'John Smith'; now type its output.
output for 'John Smith'> Smith, John
program: field(ws,1) + ', ' + field(ws,0)
preview (first 3 of 3 lines):
  → Smith, John
  → Doe, Jane
  → Turing, Alan
[1 example(s)]  :help for commands
example> :done

Data lines come from FILE; commands are read from stdin, the preview is drawn to stderr, and the accepted result is written to stdout on :done — so exform -i names.txt > clean.txt just works. Commands: IN => OUT to add an example, :N to use data line N as the next input, :undo, :list, :explain, :show N, :done, :quit.

In-line mode — sed-by-example

By default exform rewrites the whole line. --in-line instead changes only the substring that differs between your example's input and output, leaving the rest of every line untouched — the job you'd normally reach for sed to do, but without writing the pattern. exform strips the shared context from your example, learns the inner change, and generalises the matched text into a locator so it finds the same kind of token on lines it has never seen.

$ cat build.log
commit on 2021-03-05 by ana
deploy  on 1999-12-31 by ***
skipped (no date)

$ exform --in-line -e 'commit on 2021-03-05 by ana => commit on 2021/03/05 by ana' build.log
commit on 2021/03/05 by ana
deploy  on 1999/12/31 by ***
skipped (no date)

Only the date changed; everything else is byte-for-byte preserved, and lines with no match pass through unchanged. Give a second example if one is ambiguous, and use --all to rewrite every match on a line instead of just the first:

$ printf 'level=info here\nlevel=warn there\n' \
    | exform --in-line -e 'level=info x => level=INFO x' -e 'level=warn y => level=WARN y'
level=INFO here
level=WARN there

Column mode — reshape one CSV/TSV column by example

Got a CSV or TSV and only want to fix one column? --field N applies the inferred transform to column N of every row and leaves the other columns byte-for-byte intact — the awk '{ $3 = ... }' job, by example. Your examples are the cell (the column value) before=>after.

$ cat sales.csv
id,date,amount
1,2021-05-01,100
2,2022-12-31,250

$ exform --field 2 -e '2021-05-01 => 05/01/2021' -e '2022-12-31 => 12/31/2022' sales.csv
id,date,amount
1,05/01/2021,100
2,12/31/2022,250

Only column 2 changed; id, amount, and the header row are untouched (the header cell date doesn't match the rule, so it's kept). Use --field-sep to pick the delimiter — --field-sep $'\t' for TSV — and pass several columns at once with a comma list, e.g. --field 2,4. Rows shorter than the target column pass straight through, and --on-error {keep,empty,skip,fail} decides what happens to a cell the rule can't transform.

By default --field splits on the delimiter literally, which keeps untouched columns exactly as they were. For real-world CSVs with quoted fields that contain commas (or quotes, or embedded newlines), add --csv and exform parses each row as proper RFC-4180 CSV, then re-quotes the output minimally so the file stays valid:

$ cat people.csv
name,city
"Doe, John","New York, NY"
"Smith, Jane",Paris

$ exform --field 1 --csv -e 'Doe, John => John Doe' -e 'Smith, Jane => Jane Smith' people.csv
name,city
John Doe,"New York, NY"
Jane Smith,Paris

The embedded comma in "New York, NY" is preserved as a single field, and the quoting is added/removed only where needed. --csv uses only the Python standard-library csv module (no extra dependencies) and needs a single-character --field-sep.

JSONL mode — "jq by example" — --json-field

Got a stream of JSON objects, one per line (JSON Lines / .jsonl), and you only want to reshape one field's value? --json-field KEY parses each line as a JSON object and applies the inferred transform to the string at KEY, leaving every other key — and the structure and types of the record — untouched. You don't have to remember jq syntax; you show the field's value before => after.

$ cat users.jsonl
{"id": 1, "user": {"name": "John Smith"}, "active": true}
{"id": 2, "user": {"name": "Grace Hopper"}, "active": false}
{"id": 3, "user": {"name": "Alan Turing"}, "active": true}

$ exform --json-field user.name \
    -e 'John Smith  => Smith, J.' \
    -e 'Grace Hopper => Hopper, G.' users.jsonl
{"id": 1, "user": {"name": "Smith, J."}, "active": true}
{"id": 2, "user": {"name": "Hopper, G."}, "active": false}
{"id": 3, "user": {"name": "Turing, A."}, "active": true}

KEY may be a dotted path (user.name) to reach a nested value. Records where the key is missing, or whose value isn't a string, pass through unchanged; --on-error governs values the transform can't handle. Lines that aren't valid JSON objects are kept verbatim by default (use --on-error skip to drop them). It's stdlib-only — no jq, no dependencies.

Rename files by example — --rename

Bulk-renaming files is the most tedious form-filling of all — and a perfect job for programming-by-example. Pipe exform a list of paths (from ls or find), show it a couple of old-name => new-name examples, and it renames the whole batch. It's a safe dry-run by default — it prints the plan and changes nothing until you add --apply.

$ ls
IMG_0001.JPG   IMG_0002.JPG   IMG_0042.JPG   notes.txt

$ ls *.JPG | exform --rename -e 'IMG_0001.JPG => 0001.jpg' -e 'IMG_0002.JPG => 0002.jpg'
IMG_0001.JPG  ->  0001.jpg
IMG_0002.JPG  ->  0002.jpg
IMG_0042.JPG  ->  0042.jpg
exform: 3 files would be renamed (dry-run). Re-run with --apply to perform the renames.

$ ls *.JPG | exform --rename -e 'IMG_0001.JPG => 0001.jpg' -e 'IMG_0002.JPG => 0002.jpg' --apply
exform: renamed 3 files.

Two examples were enough to infer "drop the IMG_ prefix and lower-case the extension", which then applied to the unseen IMG_0042.JPG. Only each path's basename is transformed; the directory is preserved, so find . -name '*.txt' piped in renames files in place across a tree. exform refuses to clobber: if two files would map to the same name, or a target already exists, it stops without touching anything and tells you why. Use --on-error skip to leave non-matching names alone.

Emit a standalone script — --emit python

Don't want exform in your pipeline's dependencies? Infer the transform once and have exform hand you a plain, dependency-free python3 script you can commit to your repo and run anywhere. --emit python prints that script instead of transforming input; it reads lines on stdin and writes them on stdout, exactly like exform itself.

$ exform -e 'my_var_name => myVarName' -e 'home_address => homeAddress' --emit python > camelize.py
$ printf 'deep_nested_key\nlong_field_name\n' | python3 camelize.py
deepNestedKey
longFieldName

The generated code is intentionally readable (.lower(), _field(...), …), not an opaque blob, so a reviewer can see exactly what it does. Before printing anything, exform runs the generated script against your examples and refuses to emit a script that doesn't reproduce them — so you never ship a transform that silently disagrees with the examples it was derived from. Lines the transform can't handle are passed through unchanged.

Emit a portable awk one-liner — --emit awk

Reordering and re-casing delimited columns is exactly what people reach for awk for — and exactly the syntax nobody remembers. Show exform the change by example and it hands you the awk command:

$ exform -e 'john,doe,30 => DOE john' -e 'mary,smith,25 => SMITH mary' --emit awk
# text transform generated by exform (https://github.com/ingrid-owusu/exform)
# inferred program: field(,,1).upper + ' ' + field(,,0)
# reads stdin, writes stdout; lines missing an expected field are printed best-effort
awk -F, '{ print toupper($2) " " $1 }'

Paste it into any pipeline — no exform, no Python, no install. Just like the Python backend, exform runs the generated awk against your examples first and refuses (with a reason) rather than print a command that disagrees with them. The awk backend covers the common column/case/slice/zero-pad jobs; transforms it can't express faithfully in one line (regex extractors, word-folding like title/slug/camel, mixed delimiters, negative slices) are refused honestly — use --emit python for those.

Handy flags

flag meaning
-e, --example 'IN => OUT' an example (repeatable)
-E, --examples-file FILE read examples from a file, one per line
-i, --interactive live mode: type examples and watch the transform preview update across your data
--in-line sed-by-example: change only the differing substring in each line
--all in --in-line mode, rewrite every match on a line (default: first)
--field N (--col) reshape only column N (or 2,4) of a delimiter-separated row
--field-sep SEP column separator for --field (default: ,)
--json-field KEY (--jf) JSONL mode: reshape one string field (dotted path ok) of each JSON object per line
--rename bulk file-rename by example: transform each piped path's basename (dry-run by default)
--apply in --rename mode, actually perform the renames
--fill Flash Fill mode: complete a 2-column input<TAB>output table
--col-sep SEP column separator for --fill (default: TAB)
--emit python print a standalone, dependency-free script that reproduces the transform
--emit awk print a portable awk one-liner that reproduces the transform
--explain print the inferred program to stderr
-q, --quiet suppress non-fatal warnings (e.g. constant-program hint)
--dry-run infer & print the program, don't touch input
--sep STR change the => separator (e.g. --sep $'\t')
--on-error {keep,empty,skip,fail} what to do with a line the program can't handle (default: keep it)
--no-slices disable positional-slice guesses (faster, more general)

How it works

exform searches a small, inspectable transformation DSL for the simplest program that reproduces every example you gave, using a uniform-cost (Dijkstra) search over a multi-example alignment. The DSL covers the moves you actually make by hand:

  • split into fields by whitespace or a delimiter (, ; | : / @ = - _ . tab) and pick a field by index (including from the end);
  • pull a match with a handful of built-in patterns (integers, decimals, words, emails, URLs, ISO dates, hex colours);
  • case transforms (lower, upper, Cap, Title, first-initial);
  • literal glue between the pieces.

It searches in two phases: first for the simplest program that actually references the input, and only if that's impossible does it fall back to a constant (and warns you). Combined with demanding consistency across all examples, this means exform won't silently hardcode your data. The result is a program you can read (--explain) and rely on.

What it is not

exform is not a general-purpose synthesiser. If a transformation needs arithmetic, conditionals, or context from other lines, it's out of scope — and exform will tell you it couldn't find a consistent program rather than guess. Add an example, or reach for a real script.

Prior art

Programming-by-example (PBE) for strings is a well-studied idea. Microsoft's FlashFill (the research is Gulwani's PROSE framework) put it in Excel; StringSolver is a Scala implementation aimed largely at batch file renaming. exform is a deliberately small, different point in that space: a zero-dependency, pipx/uvx-installable Python CLI that behaves like an ordinary Unix filter (stdin→stdout, deterministic, offline), works line-by-line on arbitrary text, and always shows you the program it inferred so you never have to trust a black box. It is not trying to match the expressiveness of PROSE — it's trying to be the thing you actually reach for in a terminal.

Library use

from exform import synthesize

program = synthesize([("John Smith", "Smith, J."), ("Grace Hopper", "Hopper, G.")])
print(program.explain())        # field(ws,1) + ', ' + line.first + '.'
print(program.apply("Alan Turing"))  # 'Turing, A.'

Contributing

Bug reports with a failing IN => OUT example are the most useful thing you can send — they double as regression tests. See the issues tab. Licensed under the MIT License.

Every runnable example in this README is verified in CI with mdoctest, so the documented output can't silently drift from what the current code actually prints.

About

Reshape text by example: show it a couple of before⇒after lines, it infers the transform and applies it to the whole file/stream — or emits a portable awk/python one-liner. Deterministic, offline, no regex, no LLM.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages