article
وصول مفتوح
Steered Training Data Generation for Learned Semantic Type Detection
Research footprint
At a glance
- الاستشهادات
- 1
- المراجع
- 17
- Comments
- 0
Paper overview
Abstract
In this paper, we introduce STEER to adapt learned semantic type extraction approaches to a new, unseen data lake. STEER provides a data programming framework for semantic labeling which is used to generate new labeled training data with minimal overhead. At its core, STEER comes with a novel training data generation procedure called Steered-Labeling that can generate high quality training data not only for non-numeric but also for numerical columns. With this generated training data STEER is able to fine-tune existing learned semantic type extraction models. We evaluate our approach on four different data lakes and show that we can significantly improve the performance of two different types of learned models across all data lakes.
Record transparency
Publication details
- DOI
- 10.1145/3589786
- OpenAlex
- W4381328679
- Document type
- article
- Language
- EN
- Source
- Proceedings of the ACM on Management of Data
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.