Sitelet https://web.archive.org/web/20220126223314/https://github.com/kedro-org/kedro/issues/579
Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

Include yaml examples in dataset docs #579

Open
WaylonWalker opened this issue Oct 21, 2020 · 6 comments
Open

Include yaml examples in dataset docs #579

WaylonWalker opened this issue Oct 21, 2020 · 6 comments

Comments

@WaylonWalker
Copy link
Contributor

@WaylonWalker WaylonWalker commented Oct 21, 2020

Description

The most common way I look up the docs for a DataSet is to google search for things like kedro csv, which lands me in the kedro.extras.datasets.pandas.CSVDataSet docs. This is great to see the api, but it is a bit confusing that the suggested catalog method is to use yaml, but the docs are in python.

Search for the docs

image

Current page

Currently the docs look like this, and do not include good examples for creating real catalog entries with the dataset.

image

But the suggested way to add datasets to the catalog is with the yaml api, which looks like this.

image

Context

Aligning the preferred/suggested method of creating catalogs with likely entrypoints into the docs would encourage users to use that method and have less confusion for those who aren't quite sure of the difference between the python api and yaml api.

Possible Implementation

Include yaml examples in the DataSet docstrings, with backlinks to how to implement the catalog in anyway that is documented. If the python api is left in as an example there should be a link to show how to implement that example into the project.

@yetudada
Copy link
Contributor

@yetudada yetudada commented Oct 22, 2020

I've always wondered about this too. I think it would be great to see YAML equivalent examples for the datasets. I'll create hacktoberfest tickets for this.

@stale
Copy link

@stale stale bot commented Apr 12, 2021

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stale stale bot added the stale label Apr 12, 2021
@WaylonWalker
Copy link
Contributor Author

@WaylonWalker WaylonWalker commented Apr 15, 2021

Seems that there is interest in the change? How would you want this proposed change to look?

Suggestion

    Example Using Python API:
    ::

        >>> from kedro.extras.datasets.pandas import CSVDataSet
        >>> import pandas as pd
        >>>
        >>> data = pd.DataFrame({'col1': [1, 2], 'col2': [4, 5],
        >>>                      'col3': [5, 6]})
        >>>
        >>> # data_set = CSVDataSet(filepath="gcs://bucket/test.csv")
        >>> data_set = CSVDataSet(filepath="test.csv")
        >>> data_set.save(data)
        >>> reloaded = data_set.load()
        >>> assert data.equals(reloaded)
        
    Example Creating a catalog entry with the YAML API:
    ::
        
        data_set
            type: pandas.CSVDataSet
            filepath:  gcs://bucket/test.csv  
         

@stale stale bot removed the stale label Apr 15, 2021
@lorenabalan
Copy link
Member

@lorenabalan lorenabalan commented Apr 21, 2021

I like this suggestion. We welcome PRs on this. 🙂

@stale
Copy link

@stale stale bot commented Jun 20, 2021

This issue has been automatically marked as stale because it has not had recent activity. It will be closed if no further activity occurs. Thank you for your contributions.

@stale stale bot added the stale label Jun 20, 2021
@lorenabalan lorenabalan added pinned and removed stale labels Jun 21, 2021
@nicpayne713
Copy link

@nicpayne713 nicpayne713 commented Jul 11, 2021

I like this suggestion. We welcome PRs on this. 🙂

Loren,
I'm interested in participating in this issue... What is the format I should follow for PRs? One per dataset example or should I add a bunch and open one PR for them all?
If there's docs on this feel free to just "lmgtfy" your response with the link and I'll read up - it's late and I'm on mobile so I wanted to commit before losing motivation in the morning :)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Linked pull requests

Successfully merging a pull request may close this issue.

None yet
4 participants