article وصول مفتوح

Steered Training Data Generation for Learned Semantic Type Detection

  • Proceedings of the ACM on Management of Data
  • Association for Computing Machinery
Research footprint

At a glance

الاستشهادات
1
المراجع
17
Comments
0
Paper overview

Abstract

In this paper, we introduce STEER to adapt learned semantic type extraction approaches to a new, unseen data lake. STEER provides a data programming framework for semantic labeling which is used to generate new labeled training data with minimal overhead. At its core, STEER comes with a novel training data generation procedure called Steered-Labeling that can generate high quality training data not only for non-numeric but also for numerical columns. With this generated training data STEER is able to fine-tune existing learned semantic type extraction models. We evaluate our approach on four different data lakes and show that we can significantly improve the performance of two different types of learned models across all data lakes.

Record transparency

Publication details

DOI
10.1145/3589786
OpenAlex
W4381328679
Document type
article
Language
EN
Source
Proceedings of the ACM on Management of Data
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.