preprint
وصول مفتوح
PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency
Research footprint
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Paper overview
Abstract
We introduce PLaMo-100B, a large-scale language model designed for Japanese proficiency. The model was trained from scratch using 2 trillion tokens, with architecture such as QK Normalization and Z-Loss to ensure training stability during the training process. Post-training techniques, including Supervised Fine-Tuning and Direct Preference Optimization, were applied to refine the model's performance. Benchmark evaluations suggest that PLaMo-100B performs well, particularly in Japanese-specific tasks, achieving results that are competitive with frontier models like GPT-4. The base model is available at https://huggingface.co/pfnet/plamo-100b.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.2410.07563
- OpenAlex
- W4403364498
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.