New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Include yaml examples in dataset docs #579
Comments
|
I've always wondered about this too. I think it would be great to see YAML equivalent examples for the datasets. I'll create hacktoberfest tickets for this. |
|
This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions. |
|
Seems that there is interest in the change? How would you want this proposed change to look? Suggestion Example Using Python API:
::
>>> from kedro.extras.datasets.pandas import CSVDataSet
>>> import pandas as pd
>>>
>>> data = pd.DataFrame({'col1': [1, 2], 'col2': [4, 5],
>>> 'col3': [5, 6]})
>>>
>>> # data_set = CSVDataSet(filepath="gcs://bucket/test.csv")
>>> data_set = CSVDataSet(filepath="test.csv")
>>> data_set.save(data)
>>> reloaded = data_set.load()
>>> assert data.equals(reloaded)
Example Creating a catalog entry with the YAML API:
::
data_set
type: pandas.CSVDataSet
filepath: gcs://bucket/test.csv
|
|
I like this suggestion. We welcome PRs on this. |
|
This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions. |
Loren, |
Description
The most common way I look up the docs for a DataSet is to google search for things like
kedro csv, which lands me in the kedro.extras.datasets.pandas.CSVDataSet docs. This is great to see the api, but it is a bit confusing that the suggested catalog method is to use yaml, but the docs are in python.Search for the docs
Current page
Currently the docs look like this, and do not include good examples for creating real catalog entries with the dataset.
But the suggested way to add datasets to the catalog is with the yaml api, which looks like this.
Context
Aligning the preferred/suggested method of creating catalogs with likely entrypoints into the docs would encourage users to use that method and have less confusion for those who aren't quite sure of the difference between the python api and yaml api.
Possible Implementation
Include yaml examples in the DataSet docstrings, with backlinks to how to implement the catalog in anyway that is documented. If the python api is left in as an example there should be a link to show how to implement that example into the project.
The text was updated successfully, but these errors were encountered: