Large language models in healthcare: A comprehensive benchmark
a statistical and human evaluation of sixteen different LLMs applied to medical language tasks.
š§« A curated list of resources relevant to doing Biomedical Information Extraction (including BioNLP)
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
a statistical and human evaluation of sixteen different LLMs applied to medical language tasks.
a high-level review of LLM applications in medicine as of March 2024.
a review of ethical issues arising from applications of LLMs in biomedicine.
a frequently referenced but still relevant work concerning the roles, applications, and risks of language models.
An overview of how BioIE and bioinformatics workflows can be applied to questions in cardiovascular health and medicine research.
A review of clinical IE papers published as of September 2016. From Mayo Clinic group (see below).
A review of Literature Based Discovery (LBD), or the philosophy that meaningful connections may be found between seemingly unrelated scientific literature.; For some historical context on LBD, see papers by University of Chicago's Don Swanson and Neil Smalheiser, including Undiscovered Publicā¦
A review of the methods and philosophy behind mining electronic health records, including using them for adverse event detection. See Table 2 for a list of relevant papers as of mid-2017.
A 2017 review of natural language processing methods applied to information extraction in health records and social media text. An important note from this review: "One of the main challenges in the field is the availability of data that can be shared and which can be used by the community to pushā¦
This is a collection of research papers for AI-based protein design.
Led by Dr. Guergana Savova, formerly at Mayo Clinic and the Apache cTAKES project.
Based at Brown University and directed by Dr. Neil Sarkar, whose research group works on topics in clinical NLP and IE.
based at University of Colorado, Denver and led by Larry Hunter - see their GitHub repos here.
Develops improvements to biomedical literature search and curation (e.g., through PubMed), led by Dr. Zhiyong Lu.
Based at the Novo Nordisk Foundation Center for Protein Research at the University of Copenhagen, Denmark.
Based at the University of Manchester and led by Prof. Sophia Ananiadou, NaCTeM is concerned with text mining in general but has a particular focus on biomedical applications.
Several groups at Mayo Clinic have made major contributions to BioIE (for example, the Apache cTAKES platform) over the past 20 years.
A joint effort between groups at Oregon State University, Oregon Health & Science University, Lawrence Berkeley National Lab, The Jackson Laboratory, and several others, seeking to "integrate biological information using semantics, and present it in a novel way, leveraging phenotypes to bridge theā¦
Based at the University of Turku and concerned with NLP in general with a focus on BioNLP and clinical applications.
Based in the University of Texas Health Science Center at Houston, School of Biomedical Informatics and led by Dr. Hua Xu.
Based at Virginia Commonwealth University and led by Dr. Bridget McInnes.
Group led by Dr. Isaac Kohane at Harvard Medical School's Department of Biomedical Informatics (Dr. Kohane is also a steward of the n2c2 (formerly i2b2) datasets - see Datasets below).
Led by Drs. George Hripcsak and NoƩmie Elhadad.
Its subtitle is "The Journal of Biological Databases and Curation". Open access.
Nucleic Acids Research. Has a broad biomolecular focus but is particularly notable for its annual database issue.
The Journal of the American Medical Informatics Association. Concerns "articles in the areas of clinical care, clinical research, translational science, implementation science, imaging, education, consumer health, public health, and policy".
The Journal of Biomedical Informatics. Not open access by default, though it does have an open-access "X" version.
An open-access Springer Nature journal publishing "descriptions of scientifically valuable datasets, and research that advances the sharing and reuse of scientific data".
The ACM Conference on Bioinformatics, Computational Biology, and Health Informatics. Held annually since 2010.
The IEEE International Conference on Bioinformatics and Biomedicine.
The International Conference on Intelligent Systems for Molecular Biology is an annual conference hosted by the International Society for Computational Biology since 1993. Much of its focus has concerned bioinformatics and computational biology without an explicit clinical focus, though it hasā¦
The Pacific Symposium on Biocomputing.
Challenges on biomedical semantic indexing and question answering. Challenges and workshops held annually since 2013.
These workshops have been organized since 2004, with BioCreative VI happening February 2017 and the BioCreative/OHNLP Challenge held in 2018. See Datasets below.
Tasks and evaluations in computational semantic analysis. Tasks vary by year but frequently cover scientific and/or biomedical language, e.g. the SemEval-2019 Task 12 on Toponym Resolution in Scientific Papers.
Challenges for encouraging "development of software technologies to automatically extract a large variety of knowledge from eHealth documents written in the Spanish Language". Previously held as part of TASS, an annual workshop for semantic analysis in Spanish.
Held along with several other more bioinformatics-focused challenges, this challenge opened in October 2019 and focuses on using electronic health record data to predict patient mortality. Uses a synthetic data set rather than real EHR contents.
A brief introduction to bio-text mining from Cohen and Hunter. More than ten years old but still quite relevant. See also an earlier paper by the same authors.
A (non-free) volume of Methods in Molecular Biology from 2014. Chapters covers introductory principles in text mining, applications in the biological sciences, and potential for use in clinical or medical safety scenarios.
About three hours worth of video lectures on working with medical data of various types and structures, including text and image data. Appears fairly high-level and intended for beginners.
This training workshop happenened in 2013 but the slides are still online.
paper - code - Python tools primarily intended for bioinformatics and computational molecular biology purposes, but also a convenient way to obtain data, including documents/abstracts from PubMed (see Chapter 9 of the documentation).
paper - A framework for biomedical coreference resolution.
A system for building predictive medical natural language processing models. Built on the spaCy framework.
paper - A version of the spaCy framework for scientific and biomedical documents.
R utilities for accessing NCBI resources, including PubMed.
paper - code - a Python package and model (for use with spaCy) for doing NER with medication-related concepts.
Code associated with the MIMIC-III dataset (see below). Includes some helpful tutorials.
paper - code - A system for processing the text in electronic medical records. Widely used and open source.
paper - A natural language processing toolkit intended for use with the text in clinical reports. Check out their live demo first to see what it does. Usable at no cost for academic research.
A system for processing documents describing cancer presentations. Based on cTAKES (see above).
paper - A method for disease normalization, i.e., linking mentions of disease names and acronyms to unique concept identifiers. Downloadable version includes the NCBI Disease Corpus and BC5CDR (see Annotated Text Data below).
A Node.js CLI and agent skill for deterministic, daily-updated retrieval across PubMed/PMC, arXiv, and supported US policy corpora, with metadata lookup and full-document Markdown retrieval.
paper - A web platform that identifies five different types of biomedical concepts in PubMed articles and PubMed Central full texts. The full annotation sets are downloadable (see Annotated Text Data below).
A framework for running text mining tools on the newest set(s) of documents from PubMed.
paper - an IE infrastructure for electronic health records (EHR). Built on the CogStack project.
paper - Performs concept normalization (see also DNorm above). Can be trained for specific concept types and can perform NER independent of other normalization functions.
paper - a framework for IE from tables in the literature.
paper - An annotation tool with adjudication and progress tracking features.
paper - code - The brat rapid annotation tool. Supports producing text annotations visually, through the browser. Not subject specific; appropriate for many annotation projects. Visualization is based on that of the stav tool.
paper - code - An annotation tool designed to have minimal dependencies.
paper - code - A PubMed and PubMed Central-trained version of the BERT language model.
paper - A BERT model trained on >1M papers from the Semantic Scholar database.
paper - A BERT model pre-trained on PubMed text and MIMIC-III notes.
paper - A BERT model trained from scratch on PubMed, with versions trained on abstracts+full texts and on abstracts alone.
A language model available through the Flair framework and embedding method. Trained over a 5% sample of PubMed abstracts until 2015, or > 1.2 million abstracts in total.
demonstrates how text embeddings trained on biomedical or clinical text can, but don't always, perform better on biomedical natural language processing tasks. That being said, pre-trained embeddings may be appropriate for your needs, especially as training domain-specific embeddings can beā¦
paper - Qord embeddings derived from biomedical text (>10 million PubMed abstracts) using the popular word2vec tool.
paper - code - Word embeddings derived from biomedical text (>27 million PubMed titles and abstracts), including subword embedding model based on MeSH.
paper - 348,566 MEDLINE entries (title and sometimes abstract) from between 1987 and 1991. Includes MeSH labels. Primarily of historical significance.
A set of PubMed Central articles usable under licenses other than traditional copyright, though the exact licenses vary by publication and source. Articles are available as PDF and XML.
A corpus of scholarly manuscripts concerning COVID-19. Articles are primarily from PubMed Central and preprint servers, though the set also includes metadata on papers without full-text availability.
paper - A pilot dataset containing standardised information, and annotations of occurence in text, about ~5,000 known adverse reactions for 200 FDA-approved drugs.
paper - 15,000 sentences (10,000 training and 5,000 test) annotated for protein and gene names. 1,000 full text biomedical research articles annotated with protein names and Gene Ontology terms.
paper - 1,500 articles (title and abstract) published in 2014 or later, annotated for 4,409 chemicals, 5,818 diseases and 3116 chemicalādisease interactions. Requires registration.
paper - >2,400 articles annotated with chemical-protein interactions of a variety of relation types. Requires registration.
paper - 67 full-text biomedical articles annotated in a variety of ways, including for concepts and coreferences. Now on version 5, including annotations linking concepts to the MONDO disease ontology.
The Department of Biomedical Informatics (DBMI) at Harvard Medical School manages data for the National NLP Clinical Challenges and the Informatics for Integrating Biology and the Bedside challenges running since 2006. They require registration before access and use. Datasets include a variety ofā¦
paper - A corpus of 793 biomedical abstracts annotated with names of diseases and related concepts from MeSH and OMIM.
paper - A web platform that identifies five different types of biomedical concepts in PubMed articles and PubMed Central full texts. The full annotation sets are downloadable (see Annotated Text Data below).
paper - 203 ambiguous words and 37,888 automatically extracted instances of their use in biomedical research publications. Requires UTS account.
also known as CQC or the Iowa collection, these are several thousand questions posed by physicians during office visits along with the associated answers.
data from six shared tasks, though some may not be easily accessible; try the CG task set (BioNLP2013CG) for extensive entity and event annotations.
paper - a corpus of sentences from medical and biological documents, annotated for negation, speculation, and linguistic scope.
paper - a set of >6.5K biomedical relation annotations, plus labels for novel findings.
paper - 225 MEDLINE abstracts annotated for PPI.
paper - 120 full text articles annotated for PPI and genetic interactions. Used in the BioCreative V BioC task.
paper - 1,100 sentences from biomedical research abstracts annotated for relationships (including PPI), named entities, and syntactic dependencies. Additional information and download links are here.
paper - 50 scientific abstracts referenced by the Human Protein Reference Database, annotated for PPI.
paper - 486 sentences from biomedical research abstracts annotated for pairs of co-occurring chemicals, including proteins (hence, PPI annotations).
paper - 77 sentences from research articles about the bacterium Bacillus subtilis, annotated for proteināgene interactions (so, fairly close to PPI annotations). Additional information is here.
paper - A database of prevalence and co-occurrence frequencies of conditions, drugs, procedures, and patient demographics extracted from electronic health records. Does not include original record text.
paper - A database of manually curated associations between chemicals, gene products, phenotypes, diseases, and environmental exposures. Useful for assembling ontologies of the related concepts, such as types of chemicals.
paper - Deidentified health data from ~60,000 intensive care unit admissions. Requires completion of an online training course (CITI training) and acceptance of a data use agreement prior to use.
The MIMIC Chest X-Ray database. Contains more than 377,000 radiographic images and accompanying free-text radiology reports. As with MIMIC-III, requires acceptance of a data use agreement.
reference manual - A large and comprehensive collection of biomedical terminology and identifiers, as well as accompanying tools and scripts. Depending on your purposes, the single file MRCONSO.RRF may be sufficient, as this file contains unique identifiers and names for all concepts in the UMLSā¦
An update to MIMIC-III's multimodal patient data, now covering more recent years of admissions, plus a new data structure, emergency department records, and links to MIMIC-CXR images.
paper - a database of observations from more than 200 thousand intensive care unit admissions, with consistent structure. Requires registration, training course completion, and data use agreement.
paper - An ontology of human diseases. Has cross-links to MeSH, ICD, NCI Thesaurus, SNOMED, and OMIM. Public domain. Available on GitHub and on the OBO Foundry.
paper - Normalized names for clinical drugs and drug packs, with combined ingredients, strengths, and form, and assigned types from the Semantic Network (see below). Released monthly.
paper - A general English lexicon that includes many biomedical terms. Updated yearly since 1994 and still updated as of 2019. Part of UMLS but does not require UTS account to download.
paper - Mappings between >3.8 million concepts, 14 million concept names, and >200 sources of biomedical vocabulary and identifiers. It's big. It may help to prepare a subset of the Metathesaurus with the MetamorphoSys installation tool but we're still talking about ~30 Gb of disk space requiredā¦
paper - Lists of 133 semantic types and 54 semantic relationships covering biomedical concepts and vocabulary. Is the Metathesaurus too complex for your needs? Try this. Does not require UTS account to download.
code - A data model of biological entities. Provided as a YAML file.
paper - An architecture for biomedical data analysis, integration, and visualization. Conceptually based on the visual modeling language UML.
a standard for observational healthcare data.
Apache-2.0 JSON Schema (Draft 2020-12) API contract for cross-vendor somatic NGS interpretation output (Foundation Medicine, Tempus, Caris, Guardant), aligned with the HL7 FHIR Genomics IG. A standards-aligned target representation for biomedical information-extraction pipelines that parseā¦
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable teamā¦
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
analysis-tools-dev/static-analysis
āļø A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improveā¦