Sampo Pyysalo
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
The STRING database in 2021: customizable protein–protein networks, and functional characterization of user-uploaded gene/measurement sets
2020 · Nucleic Acids Research
Cellular life depends on a complex web of functional associations between biomolecules. Among these associations, protein-protein interactions are particularly important due to their versatility, specificity and adaptability. The STRING database aims to integrate all known …
-
Universal Dependencies v1: A Multilingual Treebank Collection
2016
Cross-linguistically consistent annotation is necessary for sound comparative evaluation and cross-lingual learning experiments.It is also useful for multilingual system development and comparative linguistic studies.Universal Dependencies is an open community effort to create cross-linguistically consistent treebank …
-
Attending to Characters in Neural Sequence Labeling Models
2016 · arXiv (Cornell University)
Sequence labeling architectures use word embeddings for capturing similarity, but suffer when handling previously unseen or rare words. We investigate character-level extensions to such models and propose a novel architecture for combining alternative word representations. …
-
Neural Dependency Parsing of Biomedical Text: TurkuNLP entry in the CRAFT Structural Annotation Task
2019
We present the approach taken by the TurkuNLP group in the CRAFT Structural Annotation task, a shared task on dependency parsing. Our approach builds primarily on the Turku neural parser, a native dependency parser that …
-
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection
2020 · Uppsala University Publications (Uppsala University)
Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, …
-
Exploring Cross-sentence Contexts for Named Entity Recognition with BERT
2020
Named entity recognition (NER) is frequently addressed as a sequence classification task with each input consisting of one sentence of text. It is nevertheless clear that useful information for NER is often found also elsewhere …
-
How to Train good Word Embeddings for Biomedical NLP
2016
The quality of word embeddings depends on the input corpora, model architectures, and hyper-parameter settings. Using the state-of-the-art neural embedding tool word2vec and both intrinsic and extrinsic evaluations, we present a comprehensive study of how …
-
CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies
2017
Daniel Zeman, Martin Popel, Milan Straka, Jan Hajič, Joakim Nivre, Filip Ginter, Juhani Luotolahti, Sampo Pyysalo, Slav Petrov, Martin Potthast, Francis Tyers, Elena Badmaeva, Memduh Gokirmak, Anna Nedoluzhko, Silvie Cinková, Jan Hajič jr., Jaroslava Hlaváčová, …
-
Multilingual is not enough: BERT for Finnish
2019 · arXiv (Cornell University)
Deep learning-based language models pretrained on large unannotated text corpora have been demonstrated to allow efficient transfer learning for natural language processing, with recent approaches such as the transformer-based BERT model advancing the state of …