conference-paper Open access

A Roadmap for Enriching Jupyter Notebooks Documentation with Kaggle Data

Research footprint

At a glance

Citations
0
References
4
Comments
0
Paper overview

Abstract

Recent advancements in AI and data science have led to the increased use of Jupyter notebooks. As such, various AI-Based automated tools have been also developed to automatically document notebooks. However, a key challenge is the absence of suitable datasets for training AI models. In this paper, we outline a valuable roadmap for developing a dataset of (markdown, code) pairs centered on functions in Jupyter notebooks. The roadmap encompasses four high-level steps: structural filtering, structural processing, conceptual filtering, and conceptual processing. Our proposed roadmap leads to providing a quality dataset for training AI models on Jupyter notebooks.

Record transparency

Publication details

DOI
10.1145/3644815.3644984
OpenAlex
W4399530936
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.