Florian Mai
4 papers in the PaperMetrix corpus
Papers by this author
-
Optimizer Benchmarking Needs to Account for Hyperparameter Tuning
2020 · Infoscience (Ecole Polytechnique Fédérale de Lausanne)
The performance of optimizers, particularly in deep learning, depends considerably on their chosen hyperparameter configuration. The efficacy of optimizers is often studied under near-optimal problem-specific hyperparameters, and finding these settings may be prohibitively costly for …
-
Plug and Play Autoencoders for Conditional Text Generation
2020
Text autoencoders are commonly used for conditional generation tasks such as style transfer.We propose methods which are plug and play, where any pretrained autoencoder can be used, and only require learning a mapping within the …
-
Learning to Plan for Language Modeling from Unlabeled Data
2024 · arXiv (Cornell University)
By training to predict the next token in an unlabeled corpus, large language models learn to perform many tasks without any labeled data. However, their next-token-prediction objective arguably limits their performance in scenarios that require …
-
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
2025 · arXiv (Cornell University)
High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets remains limited. Existing state-of-the-art datasets mostly rely on heuristic filtering methods, restricting both their …