article

Implementation of Federated Learning Using Probabilistic Sampling Techniques Based on Data Distribution Estimation to Solve Statistical Heterogeneity Problems

  • 한국통신학회논문지
Research footprint

At a glance

Citations
2
References
0
Comments
0
Paper overview

Abstract

연합학습 통계적 이질성이란 연합학습에 참여하는 다수의 사용자가 사용하는 디바이스, 동적 환경 및 시공간으로부터 수집된 데이터에서 IID(Independent Identically Distributed) 조건을 만족하지 못하고 불균형한 분포 특성(Non-Independent Identically Distributed)을 나타내는 것을 의미한다. 본 논문은 연합학습의 통계적 이질성 문제를 해결하기 위해 로컬 데이터 분포에 기반하여 글로벌 데이터 분포 추정하고, 확률적으로 데이터 샘플링을 수행하는 프로세스를 제안하고 직접 구현하여 성능을 비교한다. 로컬 데이터에 직접적인 접근 없이 로컬 데이터의 분포를 통해 전체 데이터의 분포를 추정하여 로컬 데이터의 분포를 조정한다. 공개된 연합학습 프레임워크에 프로세스 기능을 추가하는 형태로 구현하여 MNIST(Modified National Institute of Standards and Technology database) 데이터를 이용해 분류 모델을 학습시킨다. 일반 연합학습과 본 연구에서 제안한 샘플링 기법을 적용한 연합학습을 다양한 환경의 클라이언트에서 100라운드까지 수행한 후 성능을 비교한 결과, 평균 0.91의 Accuracy와 평균 0.98의 AUROC(Area Under the Receiver Operating Characteristic Curve)로 비슷한 수준의 성능을 보이지만 라운드당 약 9% 정도의 학습 시간 단축과 로컬 클라이언트 사이에서 발생하는 성능 이질성을 약 1.5% 감소시켰다.

Record transparency

Publication details

DOI
10.7840/kics.2021.46.11.1941
OpenAlex
W3216123935
Document type
article
Language
EN
Source
한국통신학회논문지
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.