Johannes Welbl
3 papers in the PaperMetrix corpus
Papers by this author
-
UCL Machine Reading Group: Four Factor Framework For Fact Finding (HexaF)
2018
In this paper we describe our 2 nd place FEVER shared-task system that achieved a FEVER score of 62.52% on the provisional test set (without additional human evaluation), and 65.41% on the development set. Our …
-
Training Compute-Optimal Large Language Models
2022 · arXiv (Cornell University)
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the …
-
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
2021 · arXiv (Cornell University)
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …