Sitelet https://github.com/apache/datafusion-python/pull/1789
Skip to content

fix(examples): make csv-read-options self-contained and add output to silent examples - #1789

Open
Mola-maker wants to merge 1 commit into
apache:mainfrom
Mola-maker:fix/examples-1728
Open

Mola-maker wants to merge 1 commit into
apache:mainfrom
Mola-maker:fix/examples-1728

Conversation

@Mola-maker

Copy link
Copy Markdown

Which issue does this PR close?

Partially addresses #1728: this PR covers the two example fixes described in the issue. The CI job proposed there is intentionally left out, because the issue asks for the shape of that job (including its skip matrix) to be agreed before writing it.

Rationale for this change

See #1728: no CI job runs the top-level examples/*.py scripts, so csv-read-options.py has been crashing on a missing data.csv, and nine examples produce no visible output at all, which makes them read as if they do nothing.

What changes are included in this PR?

  • examples/csv-read-options.py: the example is now self-contained. It writes its own small CSV file, plus a gzipped copy, into a temporary directory at the top and reads from those paths, instead of expecting a data.csv / data.csv.gz that does not exist in the repository.
  • Nine examples that previously ended in bare assert statements now show their result in the terminal, with the asserts kept: export.py prints each exported form, import.py shows each created DataFrame, python-udaf.py, python-udf.py and query-pyarrow-data.py show their result DataFrame, sql-using-python-udaf.py and sql-using-python-udf.py show the SQL result, sql-to-pandas.py prints the pandas DataFrame before plotting it, and substrait.py prints the logical plan recovered from the round trip.

Testing: all ten modified examples were run against the published datafusion 54.1.0 wheel. The original csv-read-options.py fails with exit code 1 (missing data.csv); the fixed version and all nine output examples exit 0 and print their results. sql-to-pandas.py was verified with a small generated yellow_tripdata_2021-01.parquet fixture (the real file is a manual download, as the issue notes) and produced both the printed aggregate table and chart.png. substrait.py was run from the repository root with the testing submodule checked out. ruff format --check passes on all ten files.

Are there any user-facing changes?

No API changes. Only the examples/ scripts change: one is fixed so it runs at all, and nine now print or show their results when run.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant