ACL Anthology
canonical archive of papers from ACL, EMNLP, NAACL, EACL, COLING, and related venues.
:book: A curated list of resources dedicated to Natural Language Processing (NLP)
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
canonical archive of papers from ACL, EMNLP, NAACL, EACL, COLING, and related venues.
tracks state-of-the-art results across common NLP tasks and datasets.
papers, benchmarks, and leaderboards for NLP tasks.
regular roundups of NLP research and trends.
the rolling review process feeding ACL-affiliated venues.
long-form essays on ML and NLP research.
illustrated summaries of recent papers.
2018 essay on the rise of pretrained language models.
2017 NLG survey.
and The Illustrated BERT, ELMo, and co. - canonical visual explanations.
Notable contributions include a tool to reconstruct long dead languages, referenced here and by taking corpora from 637 languages currently spoken in Asia and the Pacific and recreating their descendant.
Notable projects include Avenue Project, a syntax driven machine translation system for endangered languages like Quechua and Aymara and previously, Noah's Ark which created AQMAR to improve NLP tools for Arabic.
Responsible for creating BOLT ( interactive error handling for speech translation systems) and an un-named project to characterize laughter in dialogue.
Recently in the news for developing speech recognition software to create a diagnostic test or Parkinson's Disease, here.
Notable contributions include Human-Computer Cooperation or Word-by-Word Question Answering and modeling development of phonetic representations.
famous for creating the Penn Treebank and the Penn Discourse Treebank.
One of the top NLP research labs in the world, notable for creating Stanford CoreNLP and their coreference resolution system
from Google's Senior Creative Engineer explains Machine Learning for engineer's and executives alike
a16z AI playbook is a great link to forward to your managers or content for your presentations
regular roundups of NLP research and trends.
guide to managing larger linguistic annotation projects
collection of blog posts covering a wide array of NLP topics with detailed implementation
Collection of Github notebooks
NLTK Tutorials, Jupyter notebooks
An online and print book introducing NLP concepts using NLTK. The book's authors also wrote the NLTK library.
Hugging Face 🤗
Free online course covering text processing, large-scale data analysis, processing pipelines, and training neural network models for custom NLP tasks.
Beginner-friendly tutorials including getting started guides, deep learning for NLP, and visual explanations of techniques like BERT, GloVe, and TF-IDF.
and The Illustrated Transformer
by Hal Daumé III
illustrated summaries of recent papers.
CS 685, UMass Amherst CS
Lectures series from Oxford
Richard Socher and Christopher Manning's Stanford Course
Carnegie Mellon Language Technology Institute there
by Yandex Data School, covering important ideas from text embedding to machine translation including sequence modeling, language models and so on.
This covers a blend of traditional NLP topics (including regex, SVD, naive bayes, tokenization) and recent neural network approaches (including RNNs, seq2seq, GRUs, and the Transformer), as well as addressing urgent ethical issues, such as bias and disinformation. Find the Jupyter Notebooks here
Lectures go from introduction to NLP and text processing to Recurrent Neural Networks and Transformers. Material can be found here.
Lecture series from IIT Madras taking from the basics all the way to autoencoders and everything. The github notebooks for this course are also available here
4-course program covering sentiment analysis, word embeddings, RNNs, LSTMs, attention mechanisms, and Transformer models like BERT and T5 for tasks including machine translation and summarization.
end-to-end course on building language models, including data, tokenization, training, and evaluation.
seminar series with guest lectures from authors of recent transformer and NLP research.
free course on LLMs, embeddings, semantic search, and NLP applications.
hands-on NLP with Transformers, Datasets, and Tokenizers libraries.
Free beginner-friendly course covering NLP fundamentals through transformers, with Python/Jupyter notebooks.
free, by Prof. Dan Jurafsy
free, NLP notes by Dr. Jacob Eisenstein at GeorgiaTech
Brian & Delip Rao
This book serves as an introduction of text mining using the tidytext package and other tidy tools in R. Authors: Julia Silge and David Robinson.
An online and print book introducing NLP concepts using NLTK. The book's authors also wrote the NLTK library.
by Stephan Raaijmakers
by Masato Hagiwara
by Hobson Lane and Maria Dyshel
by Nicole Koenigstein
bt Tiago MOnteiro | A free FreeCodeCamp book teaching the math behind AI in plain English from an engineering point of view. It covers linear algebra, calculus, probability & statistics, and optimization theory with analogies, real-life applications, and Python code examples.
A JavaScript implementation of Twitter's text processing library
A Natural Language Processor in JS
Extensible system for analyzing and manipulating natural language
Natural Language processing in the browser
A web-based annotation tool for natural language processing (NLP)
Fast and production-ready question answering w/ DistilBERT in Node.js
Sentiment models for spacy using onnx
Adversarial attacks, adversarial training, and data augmentation in NLP
Providing a consistent API for diving into common natural language processing (NLP) tasks. Stands on the giant shoulders of Natural Language Toolkit (NLTK) and Pattern, and plays nicely with both :+1:
Higher level NLP built on spaCy
Python library to conduct unsupervised semantic modelling from plain text :+1:
Python library to produce d3 visualizations of how language differs between corpora
(archived) - A deep learning toolkit for NLP, built on MXNet/Gluon.
(archived) - An NLP research library, built on PyTorch, for developing state-of-the-art deep learning models on a wide variety of linguistic tasks.
NLP research toolkit designed to support rapid prototyping with better data loaders, word vector loaders, neural network layer representations, common NLP metrics such as BLEU
Text processing tools and wrappers (e.g. Vowpal Wabbit)
Python Natural Language Processing Library. General purpose NLP library for Python, handles some specific formats like ARPA language models, Moses phrasetables, GIZA++ alignments.
Python library for working with FoLiA, an XML format for linguistic annotation.
Python package implementing the SS3 white-box text classifier; ships with interactive visualization tools that explain predictions.
A toolkit for joint part-of-speech (POS) tagging and dependency parsing. jPTDP provides pre-trained models for 40+ languages.
a fast library for topic modelling
A production ready library for intent parsing
A library for downloading&parsing standard NLP research datasets
Word forms can accurately generate all possible forms of an English word
A multilingual and extensible document clustering pipeline
A library containing a wide variety of NLP functionality, supporting over 50 corpora.
A library for exploring the state-of-the-art deep learning topologies and techniques for NLP and NLU
A very simple framework for state-of-the-art multilingual NLP built on PyTorch. Includes BERT, ELMo and Flair embeddings.
Simple, Keras-powered multilingual NLP framework, allows you to build your models in 5 minutes for named entity recognition (NER), part-of-speech tagging (PoS) and text classification tasks. Includes BERT and word2vec embedding.
Fast & easy transfer learning for NLP. Harvesting language models for the industry. Focus on Question Answering.
End-to-end Python framework for building natural language search interfaces to data. Leverages Transformers and the State-of-the-Art of NLP. Supports DPR, Elasticsearch, HuggingFace’s Modelhub, and much more!
a DSL, loosely based on RUTA on Apache UIMA. Allows to define language patterns (rule-based NLP) which are then translated into spaCy, or if you prefer less features and lightweight - regex patterns.
Natural Language Processing for TensorFlow 2.0 and PyTorch.
Tokenizers optimized for Research and Production.
Facebook AI Research implementations of SOTA seq2seq models in Pytorch.
Hierarchical Topic Modeling with Minimal Domain Knowledge
Neural Machine Translation (NMT) toolkit that powers Amazon Translate.
A deep learning-based translation library for 50 languages, built on transformers and Facebook's mBART Large.
Evaluation of NLP model outputs offering various automated metrics.
Unicode-aware regular-expression based tokenizer for various languages. Python binding to C++ library, supports FoLiA format.
Human annotation tool for multilingual NLP tasks, such as machine translation.
Stanford NLP's Python toolkit for tokenization, POS, lemma, dependency parsing, and NER across 70+ languages.
sentence/document embeddings, semantic search, and re-ranking; current standard for retrieval-style NLP.
open-source data annotation and feedback collection platform for LLM and NLP datasets.
standardized loaders and processing for thousands of NLP datasets.
reference implementations for NLP metrics.
reproducible BLEU/chrF/TER scoring for machine translation.
learned MT metrics, current de-facto standard.
60+ test types for NLP model robustness, bias, and fairness.
High-accuracy, rule-based sentence boundary detector (SBD). Drop-in pysbd adapter, streaming APIs, CLI, and a spaCy component across 39+ languages.
A neural network library for building instance-dependent NLP models with padding-free dynamic batching.
C, C++, and Python tools for named entity recognition and relation extraction
Open source implementation of Conditional Random Fields (CRFs) for segmenting/labeling sequential data & other Natural Language Processing tasks.
CRFsuite is an implementation of Conditional Random Fields (CRFs) for labeling sequential data.
BLLIP Natural Language Parser (also known as the Charniak-Johnson parser)
C++ library, command line tools, and Python binding for extracting and working with basic linguistic constructions such as n-grams and skipgrams in a quick and memory-efficient way.
Unicode-aware regular-expression based tokenizer for various languages. Tool and C++ library. Supports FoLiA format.
C++ library for the FoLiA format
Memory-based NLP suite developed for Dutch: PoS tagger, lemmatiser, dependency parser, NER, shallow parser, morphological analyzer.
ModErn Text Analysis: a C++ data sciences toolkit for mining big text data.
a library from Facebook for creating embeddings of word-level, paragraph-level, document-level and for text classification
adaptive probabilistic top-down and bottom-up parsers
is an open-source library for a machine learning based toolkit used in the processing of natural language text. It features an API for use cases like Named Entity Recognition, Sentence Detection, POS(Part-Of-Speech) tagging, Tokenization Feature extraction, Chunking, Parsing, and Coreference…
Web-Scale Open Information Extraction
An efficient and flexible token-based regular expression language and engine.
Core libraries developed in the U of Illinois' Cognitive Computation Group.
MAchine Learning for LanguagE Toolkit - package for statistical natural language processing, document classification, clustering, topic modeling, information extraction, and other machine learning applications to text.
A robust POS tagging toolkit available (in both Java & Python) together with pre-trained models for 40+ languages.
A language detection library for Kotlin and Java, suitable for long and short text alike
an index-based text data generator written in Kotlin
Library for developing NLP systems, including built in modules like SRL, POS, etc.
Toolkit with state-of-the-art automatic term recognition methods.
Implementation of topic modeling based on regularized multilingual PLSA.
Scala interface to word2vec model; includes operations on vectors like word-distance and word-analogy.
Epic is a high performance statistical parser written in Scala, along with a framework for building complex structured prediction models.
Spark NLP is a natural language processing library built on top of Apache Spark ML that provides simple, performant & accurate NLP annotations for machine learning pipelines that scale easily in a distributed environment.
Fast vectorization, topic modeling, distances and GloVe word embeddings in R.
An R package for creating and exploring word2vec and other word embedding models
R package to interface with the Java machine learning tool MALLET
Creates d3 visualizations for browsing topic models of text in a web browser.
R package for exploring topic models of text.
Sentiment Classification using Word Sense Disambiguation and WordNet Reader
Japanese Natural Langauge Processing Libraries, with Japanese sentiment classification
An R package for dynamic exploration of text collections
Text mining using tidy tools
R wrapper to spaCy NLP
Natural Language Processing in Clojure (opennlp)
Rails-like inflection library for Clojure and ClojureScript
A library to parse natural language in Clojure and ClojureScript
Text processing library supporting tokenization, part-of-speech tagging, and named-entity extraction.
Converts numbers to Russian words with correct grammatical gender and noun declension.
A curated list of resources dedicated to Natural Language Processing and text processing for Ruby.
Natural language recognition library based on trigrams
(archived — Snips was discontinued) - A production ready library for intent parsing
NLP++ Language Extension for VSCode
NLP++ engine to run NLP++ code on Linux including a full English parser
Homepage for the NLP++ Language
Wiki entry for the NLP++ language
A variety of loaders for various NLP corpora
A package for working with human languages
Julia package for text analysis
Neural Network based models for Natural Language Processing
High performance tokenizers for natural language processing and other related tasks
Julia interface to word2vec
Natural Language Interface for apps and devices
API and Github demo
NLP and ML suite covers most common tasks like NER, tagging, and sentiment analysis
Syntax Analysis, NER, Sentiment Analysis, and Content tagging in atleast 9 languages include English and Chinese (Simplified and Traditional).
High level Text Analysis API Service ranging from Sentiment Analysis to Intent Analysis
The Text Analytics API is a suite of text analytics web services built with best-in-class Microsoft machine learning algorithms.
Natural Language Processing in the Browser with sentiment analysis, named entity extraction, POS tagging, word frequencies, topic modeling, word clouds, and more
SpaCy NLP models (custom and pre-trained ones) served through a RESTful API for named entity recognition (NER), POS tagging, and more.
Unified and free NLP APIs that perform actions such as speech tagging, text rephrasing, language translation/detection, and sentence parsing
General Architecture and Text Engineering is 15+ years old, free and open source
is free and open source, web-based raw text annotation tool
brat rapid annotation tool is an online environment for collaborative text annotation
doccano is free, open-source, and provides annotation features for text classification, sequence labeling and sequence to sequence
A semantic annotation platform offering intelligent assistance and knowledge management
is an annotation tool powered by active learning, costs $
Hosted and managed text annotation tool for teams, costs $
open source local or online tool for discourse tree annotations
open source server annotation tool with GitHub version control and validation for XML data and collaborative spreadsheet grids
support various NLP tasks for individual or teams, freemium based
team-first hosted and on-prem text, image and PDF annotation tool powered by active learning, freemium based, costs $
Easy-to-use text annotation tool for teams with most comprehensive auto-annotation features. Supports NER, relations and document classification as well as OCR annotation for invoice labeling, costs $
Shoonya is free and open source data annotation platform with wide varities of organization and workspace level management system. Shoonya is data agnostic, can be used by teams to annotate data with various level of verification stages at scale.
Free End-to-End No-Code platform for text annotation and DL model training/tuning. Out-of-the-box support for Named Entity Recognition, Classification, Relation extraction and Assertion Status Spark NLP models. Unlimited support for users, teams, projects, documents. Not FOSS.
FLAT is a web-based linguistic annotation environment based around the FoLiA format, a rich XML-based format for linguistic annotation. Free and open source.
open-source data annotation and feedback collection platform for LLM and NLP datasets.
open-core multi-modal labeling platform; widely used for NLP labeling.
Free, open-source annotation tool covering 21+ task types (classification, span, coreference, entity linking, agent trace evaluation) with built-in MACE quality control, attention checks, AI-assisted labeling, and 300+ example tasks.
implementation - explainer blog
explainer blog
implementation; subword n-grams handle OOV well, still useful for low-resource languages.
word sense disambiguation.
deep contextualized word representations.
contextualized vectors learned from MT.
language-model fine-tuning for text classification.
sentence representations from NLI.
language-agnostic subword tokenization.
and Unigram LM - the two dominant subword schemes.
Stanford NLP's Python toolkit for tokenization, POS, lemma, dependency parsing, and NER across 70+ languages.
tokenization, tagging, lemmatization, parsing for Universal Dependencies.
unsupervised morphological segmentation. Tokenizer research and architecture (also see Language Models):
Tokenizers optimized for Research and Production.
tokenizer-free byte-level model.
tokenization-free encoder operating on Unicode characters.
tokenizer fairness across languages.
(Meta, 2024) - dynamic byte-level patching that matches BPE-tokenized models at scale; revives the tokenizer-free direction.
(2025) - superword tokenization that improves on BPE for downstream tasks.
(ICML 2025) - decouples input and output vocabularies; shows a log-linear relationship between input vocabulary size and training loss, scaling vocabulary independently of model size.
(ICLR 2025) - first formal unified framework for tokenizer models using stochastic-map category theory; establishes conditions for statistical consistency.
(2025) - quantifies how tokenization fertility predicts model accuracy across languages, exposing structural cost penalties for morphologically complex and low-resource languages.
(2026) - post-hoc vocabulary additions that coalesce multi-token character sequences for low-resource languages, reducing inference cost without retraining.
cross-linguistically consistent treebanks, 100+ languages.
foundational neural parsing architecture.
light-weight transformer-based multilingual NLP toolkit.
strong neural constituency parser.
canonical English NER benchmark.
BiLSTM-CRF, the long-time go-to NER architecture.
production-ready.
instruction-tuned LM for open-set NER across languages.
(2023) - small, generalist NER model that handles arbitrary entity types at inference.
guideline-following information extraction with LMs.
end-to-end relation extraction as seq2seq.
LLMs for named entity recognition.
(2024) - cost-quality tradeoffs.
(2026) - eight open LLMs across four NER benchmarks; PEFT with structured outputs matches encoder-based NER.
foundation for modern neural coreference.
span-based pretraining; strong coreference baseline.
higher-order inference coreference.
(2024) - efficient coreference matching the best larger systems.
linguistically-motivated category-based coreference scoring. LLM-based:
prompting and fine-tuning for coreference.
(2025) - 9 systems across 4 LLM-based and 5 traditional approaches; traditional methods still lead but LLMs are closing the gap.
strong, fast linear baseline.
canonical fine-grained sentiment dataset.
few-shot text classification without prompts.
fast few-shot for many-class settings.
current encoder-fine-tuning baseline.
Python package implementing the SS3 white-box text classifier; ships with interactive visualization tools that explain predictions.
using LLMs for text classification labeling, with caveats.
foundational topic model.
Python library to conduct unsupervised semantic modelling from plain text :+1:
a fast library for topic modelling
clustering-based topic modeling on top of contextual embeddings; common modern default.
jointly learns topic and document vectors.
Hierarchical Topic Modeling with Minimal Domain Knowledge
extractive graph-based summarization.
foundational neural abstractive summarization.
gap-sentences pretraining for summarization.
widely used denoising seq2seq baseline.
and SCROLLS - long-document summarization benchmarks. LLM-based:
LLMs vs fine-tuned summarizers.
structured prompting for summarization.
(2025) - explicit reasoning improves fluency but hurts factual grounding; longer reasoning budgets can harm faithfulness.
transformer; reset the field.
efficient C++ NMT framework.
PyTorch sequence modeling toolkit.
MT for 200 languages.
400+ language MT.
speech and text MT, 100+ languages.
learned MT metrics, current de-facto standard.
reproducible BLEU/chrF/TER scoring for machine translation.
similarity-based generation metric.
LLMs as machine translation systems.
(2024) - LLMs for context-aware translation.
quality comparison on professional MT.
(2025) - benchmarks sub-10B open LLMs on 28-language MT; matches GPT-4-turbo and Google Translate.
(2025) - survey of how instruction-following, in-context learning, and preference alignment have restructured MT methodology.
extractive reading comprehension.
real-user questions over Wikipedia.
multi-hop reasoning.
distantly-supervised QA.
open-domain QA over Wikipedia.
multi-paragraph reading comprehension.
and FiD - retrieve-then-read; the standard pre-LLM open-domain QA pipeline.
retrieval-augmented LM for few-shot QA.
(2023) - retrieval, generation, and self-critique.
general AI assistant benchmark including multi-step QA.
schema-free open information extraction.
end-to-end relation extraction as seq2seq.
document-level relation extraction benchmark.
(2025) - generative LLMs with RAG and self-correction surpass encoder-decoder BERT-style models on SRL in English and Chinese.
(2025) - decoder-only LLMs with a novel error-rate adaptation schedule set new SOTA on BEA-test grammatical error correction.
and FiD - retrieve-then-read; the standard pre-LLM open-domain QA pipeline.
and ColBERTv2 - late-interaction retrieval; strong on out-of-domain.
and E5-Mistral - widely used dense embedding families.
and BGE-M3 (2024) - multilingual, multi-functionality embeddings; top of MTEB across languages.
(2024) - fully open, reproducible embedding model.
nested embeddings supporting variable dimensionality at inference.
(2024) - unified generation and embedding from one model.
the original retrieval-augmented framework; foundation for modern QA pipelines.
(2025) - Gemini-derived dense embeddings; SOTA on MMTEB across 250+ languages and on cross-lingual retrieval (XOR-Retrieve, XTREME-UP).
(2025) - decoder-based embedding series (0.6B-8B) built on Qwen3; #1 on MTEB Multilingual and MTEB Code, surpassing prior proprietary models.
(2025) - first reranking model trained with test-time compute via DeepSeek-R1 reasoning-trace distillation; SOTA on instruction-following and OOD retrieval.
(2025) - embedding model for reasoning-intensive retrieval with ReMixer data synthesis and Redapter adaptive training; record nDCG@10 of 38.1 on BRIGHT.
(2026) - extends late-interaction retrieval by integrating query and document attention weights into ColBERT scoring; improves recall on MS-MARCO, BEIR, and LoTTE. Embedding and retrieval benchmarks:
(2025) - community expansion of MTEB to 500+ tasks across 250+ languages.
unified speech and text translation.
(NVIDIA, 2024) - top open multilingual ASR model.
industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.
foundational self-supervised speech pretraining.
the central index for modern NLP datasets, with versioned, streamable loaders.
large collection of NLP datasets.
data repository for pretrained NLP models and NLP corpora.
825 GiB diverse text corpus.
(2023-2024) - reproductions of LLaMA pretraining data; V2 is 30T tokens with quality signals.
(AI2, 2023-2024) - 3T-token open pretraining corpus with documented filtering pipeline.
(2024) - 15T-token cleaned web corpus; FineWeb-Edu filters for educational quality.
6.3T tokens across 167 languages.
(2024) - 2T-token open-license multilingual corpus.
cross-linguistically consistent treebanks, 100+ languages.
(2024) - open instruction-tuning data behind Tülu 3.
tiny NLP multi-lingual QA datasets and library to generate your own synthetic copies.
tokenization, tagging, lemmatization, parsing for Universal Dependencies.
Natural Language Processing Pipeline - Sentence Splitting, Tokenization, Lemmatization, Part-of-speech Tagging and Dependency Parsing. New platform, written in Python with Dynet 2.0. Offers standalone (CLI/Python bindings) and server functionality (REST API).
is an NLP library mostly for many endangered Uralic languages such as Sami languages, Mordvin languages, Mari languages, Komi languages and so on. Also some non-endangered languages are supported such as Finnish together with non-Uralic languages such as Swedish and Arabic. UralicNLP can do…
bidirectional transformer pretraining; foundation for most encoder-based NLP work since 2018. Read online with section navigation and the ACL source attached.
robustly optimized BERT pretraining; common encoder baseline.
current encoder-fine-tuning baseline.
replaced-token-detection pretraining, sample-efficient.
(2024) - modernized encoder with rotary embeddings, FlashAttention, 8K context; current go-to encoder for classification, NER, retrieval.
(2025) - 250M-parameter encoder integrating modern architecture improvements (RoPE, 4K context, optimized depth-to-width); state of the art on MTEB, surpasses ModernBERT and RoBERTa-large under identical fine-tuning.
and FLAN-T5 - text-to-text framing for NLP tasks; strong instruction-tuned encoder-decoder baselines.
widely used denoising seq2seq baseline.
(Meta, 2024-2025) - widely adopted open-weight family; default base for fine-tuning across NLP tasks.
(Alibaba, 2024-2025) - strong multilingual coverage, especially Chinese; often top open model on multilingual benchmarks.
(2024) - efficient MoE pretraining; competitive open base model.
(AI2, 2025) - fully open: weights, training data, code; reproducibility benchmark.
(Google, 2024-2025) - open small/mid-size models with strong NLP-task performance.
efficient dense and sparse-MoE open models.
encoder vs decoder vs encoder-decoder for NLP transfer.
cross-lingual masked LM trained on CommonCrawl, 100 languages.
multilingual T5 covering 101 languages.
176B-parameter open multilingual LM, 46 natural languages.
(Cohere For AI, 2024) - massively multilingual instruction-tuned models covering 23-101 languages.
encoder for 500+ languages, focus on low-resource.
MT for 200 languages.
400+ language MT.
speech and text MT, 100+ languages.
(2024-2025) - LMs targeting Southeast Asian languages.
(2025) - open multilingual LLMs (9B and 83B) covering the top 25 languages by speaker population (~90% of global speakers); surpasses comparably-sized open multilingual models on XCOPA, XNLI, MGSM, FLORES-200.
(Princeton/Mila, 2025) - Llama-3.1-8B adapted for low-resource African languages via the curated WURA corpus; SOTA open-source results on IrokoBench and AfriQA.
(McGill, 2026) - suite of open LLMs (4B-14B) continued-pretrained on 26B tokens across 20 African languages with a comprehensive empirical study of data mixing.
(Google, 2026) - open translation-specialized models built on Gemma 3, covering 55 language pairs via SFT and RL with quality-reward models.
(Xiaomi, 2026) - open multilingual MT scaled across 46 languages, matching commercial systems like Google Translate and Gemini 3 Pro.
and SuperGLUE - English NLU benchmarks.
and XGLUE - cross-lingual NLU.
cross-lingual natural language inference, 15 languages.
MT evaluation across 200 languages.
Massive Text Embedding Benchmark; standard for sentence/document encoders.
heterogeneous IR benchmark for retrieval models.
holistic evaluation across NLP tasks, accuracy and beyond.
200+ tasks probing language model capabilities.
multitask knowledge evaluation across 57 subjects.
(2024) - harder, more discriminative successor to MMLU.
graduate-level Q&A, "Google-proof" reasoning evaluation.
(2026) - scientific reasoning benchmark for evidence-grounded critique, overclaim detection, missing-evidence refusal, and calibration.
verifiable instruction-following evaluation.
human-preference ELO leaderboard for chat models.
(2024) - contamination-resistant benchmark with monthly refresh.
unified framework for LM benchmark evaluation.
(2025) - multilingual extension of MMLU-Pro to 29 typologically diverse languages; reveals up to 24.3% performance gap between high- and low-resource languages.
(2025) - multi-turn conversational benchmark exposing simultaneous instruction-following and in-context-reasoning failures; all tested frontier models score below 50%.
(2025) - unified RAG evaluation: 824 multi-hop questions requiring factuality, retrieval accuracy, and cross-document reasoning together.
retrieval probe for long-context windows.
(2024) - synthetic long-context tasks beyond simple retrieval.
bilingual long-context benchmark across NLP tasks.
(2025) - 503 expert-crafted multiple-choice questions spanning 8K-2M-word contexts with deep multi-hop reasoning; humans score 53.7% under time pressure.
(2025) - extends needle-in-haystack with multi-needle and nested configurations; shows RAG mitigates lost-in-the-middle for smaller LLMs but degrades reasoning models.
foundational result; intermediate reasoning steps improve performance.
majority vote over sampled CoT chains.
search over reasoning trees.
and Reflexion - self-correction at inference time.
chain-of-thought for NLP reasoning tasks.
process-supervised reward models for reasoning.
(2025) - open reasoning model trained with pure RL; replicated o1-style behavior in the open.
(2024-2025) - test-time-compute reasoning systems.
(2024) - systematic study of inference-time compute tradeoffs.
(2025) - small open reasoning recipe via budget-forcing.
(2025) - long-context RL with policy optimization (no MCTS, no PRM) reaching o1-level performance; introduces long-CoT distillation into short-CoT models.
(2025) - small policy model paired with a process preference model trained via MCTS rollouts; enables small LMs to bootstrap reasoning without distilling from larger models.
(2025) - open GRPO-based RL training system with four key improvements (decoupled clipping, dynamic sampling, token-level loss, entropy bonus); reproduces and surpasses DeepSeek-R1-Zero-level reasoning.
(2025) - value-model-based RL with length-adaptive GAE and token-level clipping; surpasses value-free GRPO methods on AIME 2024 with stable training.
(2025) - generative process reward models that produce chain-of-thought verification per step, matching discriminative PRMs with 1% of the supervision labels.
(2025) - 1000+ controlled experiments on data recipes for open reasoning models; SOTA on AIME 2025 matching closed distillation baselines.
and Mamba-2 - selective state-space models, linear-time long-context alternative to attention.
RNN-transformer hybrid scaling to large parameter counts.
(2024) - hybrid Mamba-Transformer-MoE architecture.
and YaRN - rotary position embeddings and context-length extension.
extending context windows with minimal fine-tuning.
long-context degradation patterns in NLP tasks.
(2024) - tradeoffs for QA over long inputs.
(2025) - neural long-term memory module that learns to memorize historical context at test time; scales beyond 2M tokens, outperforms transformers and modern linear-recurrent models on language modeling and reasoning.
(2025) - 456B-parameter hybrid combining lightning (linear) attention with sparse softmax attention; matches GPT-4o-level NLP performance at up to 4M-token inference contexts.
(2025) - trainable sparse attention combining coarse-grained compression with fine-grained selection; large speedups at 64K with no NLP-benchmark degradation.
(2025) - identifies undertraining of high-frequency RoPE dimensions and applies evolutionary-search rescaling; extends LLaMA3-8B to 128K with 80x fewer training tokens than Meta's recipe.
(2025) - first comprehensive memory and speed analysis of transformer, SSM, and hybrid models up to 220K tokens; SSMs are up to 4x faster, hybrids balance recall and efficiency.
taxonomy and mitigation strategies.
benchmark for truthfulness in question answering.
fine-grained factual precision in long-form generation.
(2024) - long-form factuality benchmark and search-augmented evaluator.
sampling-based hallucination detection.
reference-free evaluation for RAG and QA pipelines.
(2024) - attention-pattern-based hallucination detection in long-context generation.
(2024) - calibration analysis under format effects.
(2025) - hallucination benchmark with extrinsic/intrinsic taxonomy and dynamic test-set regeneration to resist data leakage.
(2025) - claim-level calibration analysis for long-form generation; models are substantially worse-calibrated on extended outputs than on single claims.
(2025) - faithfulness-aware uncertainty quantification for RAG fact-checking; formally separates faithfulness from factuality.
(2025) - multilingual claim-hallucination benchmark across English, French, Spanish, German with token-level logits released for principled UQ evaluation.
(2026) - hard multi-turn hallucination benchmark for citation-required responses; ~30% hallucination rates persist even with web search.
(2026) - trains models to reason about claim-level uncertainty before generating; large gains on biography factuality and FactBench AUROC.
what BERT learns about language.
methodology, limitations, alternatives.
causal tracing of factual recall.
structural probing for linguistic knowledge.
foundation for the sparse-feature view of transformer representations.
(Anthropic, 2024) - sparse autoencoders extracting interpretable features from production-scale LMs.
SAE methodology for LM interpretability.
open platform browsing SAE features across models.
(2023) - identifying training examples driving model behavior.
(Anthropic, 2025) - introduces cross-layer transcoders and attribution graphs to construct an interpretable replacement model; enables prompt-level circuit tracing of feature-to-feature causal interactions.
(Anthropic, 2025) - applies attribution graphs to Claude 3.5 Haiku across multi-hop reasoning, rhyme planning, and jailbreak case studies.
(2025) - shows transcoders (reconstructing layer outputs from inputs) yield more interpretable features than SAEs; introduces skip transcoders.
(EMNLP 2025) - reference survey of SAE architectures, training strategies, feature explanation, and evaluation.
(2026) - identifies circuits at the per-prompt level (rather than per-task); reveals mechanism clustering by prompt family.
and MiniLM - distilled encoders for production NLP.
(Microsoft, 2024) - small models trained on curated data, competitive with much larger ones on NLP benchmarks.
(HuggingFace, 2025) - fully open small-LM family with reproducible training data.
(HuggingFace, 2025) - 3B fully open decoder pretrained on 11.2T tokens with NoPE and YaRN for 128K context; competitive with 4B-class models.
(Google, 2025) - 1B-27B open models with high local-to-global attention ratio to keep KV-cache tractable at 128K context.
(Alibaba, 2025) - dense and MoE models 0.6B-235B with unified thinking/non-thinking modes; the 30B-A3B MoE matches larger dense models while activating only 3B parameters.
(Apple, 2025) - on-device 3B model using KV-cache sharing and 2-bit QAT for 37.5% cache memory reduction without accuracy loss.
sentence and paragraph embeddings via Siamese BERT.
few-shot text classification without prompts.
fast few-shot for many-class settings.
, BGE, and Stella - compact text embedding models near the top of MTEB.
post-training quantization for transformers.
activation-aware weight quantization.
(ICML 2025) - sensitivity-aware layer-wise mixed-precision KV-cache quantization; up to 21% throughput improvement over uniform KV8.
portable quantized inference.
HF production serving for LMs.
and QLoRA - low-rank adapters and quantized fine-tuning; the standard for adapting LMs to NLP tasks on modest hardware.
(2024) - weight-decomposed low-rank adaptation.
finetuned language models as zero-shot learners.
training LMs to follow instructions with human feedback.
bootstrapping instruction data from LMs.
1600+ NLP tasks with instructions.
training LMs with AI-generated feedback against a written constitution.
simpler alternative to RLHF; widely adopted.
(AI2, 2024) - fully open post-training recipe with state-of-the-art results among open models.
"less is more for alignment"; small high-quality SFT data goes a long way.
(2024-2025) - synthesizes high-quality instruction-response pairs by prompting aligned LMs with nothing; SFT on the filtered subset matches official Llama-3-Instruct.
measuring stereotypical bias in pretrained LMs.
social bias measurement in masked LMs.
gender bias in coreference resolution.
bias measurement across many demographic axes.
toxicity in LM generation.
models tailoring answers to user beliefs.
(Anthropic, 2024) - models strategically complying during training.
(2024) - open safety moderation model and benchmark.
(2025) - finetuning on a narrow task (insecure code) unexpectedly produces broad alignment failures across unrelated domains.
(2025) - multilingual (Chinese/English) safety benchmark of 4000+ multi-turn dialogues across 22 scenarios and 7 jailbreak strategies.
(2025) - modular jailbreak evaluation framework integrating 19 attacks, 29 defenses, and 19 evaluation methods across 14 models and 12 risk categories.
(2026) - multilingual safety benchmark across 12 Indic languages; reveals 12.8% cross-language agreement, with over-refusal in low-resource scripts.
(2026) - alignment faking occurs in models as small as 7B in 37% of cases when policy conflicts with internalized values; steering-vector mitigation reduces it 94%.
Python toolkit for Arabic NLP including dialect ID, morphology, NER.
Go package for Arabic text processing.
JavaScript Arabic stemmer.
Python library for Arabic.
trainable segmenter for Arabic, Hebrew, and Coptic.
QCRI segmentation, POS tagging, and NER for Arabic.
Arabic BERT family.
BERT models for MSA, dialectal, and Classical Arabic.
efficient Arabic pretraining (released alongside AraBERT).
(2023-2024) - bilingual Arabic-English open LM family.
(SDAIA, 2024) - Arabic-first foundation models.
largest available multi-domain Arabic sentiment analysis resources.
large Arabic book reviews dataset.
aggregated Arabic stopwords.
(2024) - Arabic MMLU benchmark.
whole-word masking BERT for Chinese.
improved Chinese BERT with MLM-as-correction pretraining.
Alibaba's open Chinese-strong LM family.
Tsinghua's bilingual Chinese-English LMs.
open Chinese LM.
01.AI's bilingual open LMs.
efficient open MoE model with strong Chinese.
large collection of Chinese NLP tools and resources.
NLP resources in Danish.
curated list of resources for Danish language technology.
Python binding to Frog, an NLP suite for Dutch (POS tagging, lemmatization, dependency parsing, NER).
Dutch surface realiser for natural language generation, based on the SimpleNLG implementation.
dependency parser for Dutch (also does POS tagging and lemmatization).
Dutch speech-recognition models based on Kaldi.
industrial-strength NLP with a Dutch pipeline.
curated list of open-access, open-source, and off-the-shelf resources and tools developed with a focus on German.
A multi-representational multi-layered treebank for Hindi and Urdu
A smaller part of the above-mentioned treebank.
60k Words POS Tagged, Bangla, Hindi, Marathi, Telugu
~1k Samples, 3 polarity classes
4.3k Samples, 14 classes
5.4k Samples, 12 Domains, 4k aspect terms, aspect and sentence level polarity in 4 classes
5.5k Samples, 2 Domains, 10 aspect terms
2k Samples, 3 polarity labels
Twitter and Facebook labelled sentiment samples in Hindi, Bengali, Tamil, Telugu.
Sentiwordnet, parallel labelled corpora, sense-annotated corpora, and Marathi polarity-labelled corpus.
and nlp-for-hindi ULMFIT style languge model
2k Samples, 3 polarity labels
Trained on Sanskrit Wikipedia and OSCAR corpus
deep morphological parser for Hindi and Urdu.
tokenization, transliteration, MT helpers across 18 Indic languages.
dependency parsing and POS tagging for Kannada, Hindi, and Telugu.
NLP toolkit for Indic languages on PyTorch/Fastai.
tools, datasets, and models across 22 Indic languages.
(2022-2024) - multilingual BERT for 23 Indic languages.
(2023-2024) - high-quality MT for 22 Indic languages.
(Sarvam AI, 2023) - bilingual Hindi-English LLaMA continuation.
(2024) - instruction-tuned Hindi LLM.
(2024) - multilingual LM trained from scratch on 10 Indic languages.
(2024) - Indic-focused foundation models.
natural language toolkit for Indonesian.
trained on Wikipedia.
Python stemmer for Bahasa Indonesia based on the Sastrawi stemming algorithm.
pretrained Indonesian LM with the IndoNLU benchmark suite.
alternative IndoBERT with the IndoLEM benchmark.
(2023-2024) - large-scale community datasets and Cendol instruction-tuned LMs for Indonesian and regional languages.
open Southeast-Asian LMs covering Indonesian.
(2024) - Singapore AI's open Southeast-Asian LM with strong Indonesian.
39K sentences and 900K word tokens
10K sentences and 250K word tokens.
and Universal Dependencies-Indonesian
text summarization and classification.
large, free, semantic dictionary.
A multilingual and multimodal data hub providing standardized datasets and benchmarks for Southeast Asian NLP (EMNLP 2024).
Python package for Korean natural language processing.
C++ library for Korean NLP.
Scala library for Korean NLP.
R package for Korean NLP.
Korean sentence splitter.
fast Korean morphological analyzer.
browser-native Korean morphological analyzer running fully client-side via WebAssembly (1MB model, offline, MIT).
Korean BERT from SKT.
models trained on the KLUE benchmark.
open Korean LMs.
(LG, 2024) - bilingual Korean-English open LM family.
Naver's Korean foundation model.
corpus from the Korea Advanced Institute of Science and Technology in Korean.
dataset in Korean from a major South Korean newspaper.
chatbot data in Korean.
expired petition data from the Blue House National Petition Site.
NMT dataset for Korean to French and Korean to English.
Korean SQuAD dataset (v1.0 and v2.1) with Wiki HTML source.
Persian BERT.
(2023-2024) - Persian instruction-tuned LM.
(Part AI, 2024) - Llama-3-based Persian instruction model.
tagged corpus suitable for Persian (Farsi) NLP research, ~2.6M manually tagged words across 40 POS tags.
large freely available Persian corpus, 2.7M tokens annotated with 31 POS tags.
LSCP: 120M sentences from 27M casual Persian tweets with dependency, POS, and sentiment annotations.
250K tokens, 7,682 sentences with NER tags in IOB format.
~25M tokens, ~1M Persian sentences from Persian Wikipedia Corpus.
first Persian dataset for relation extraction (translated SemEval-2010 Task 8).
29,982 annotated sentences covering most verbs of the Persian valency lexicon.
dependency-based syntactically annotated corpus.
standard reliable Persian text collection used at CLEF 2008-2009.
curated list of Portuguese NLP resources and tools.
Python library to detect, censor, and clean profanity, hate speech, and bullying in Spanish, with data from 21 Spanish-speaking countries.
BERT for Spanish.
Spanish RoBERTa trained on the Spanish National Library corpus.
(2024) - open foundation LM for Basque, also covers Spanish.
(BSC, 2024) - multilingual LM with strong Spanish coverage from the Barcelona Supercomputing Center.
(2024) - Spanish-instruction-tuned open model.
Thai NLP in Python.
character cluster library in Java.
word segmentation with deep learning in TensorFlow.
tokenization and POS tagging.
word segmentation and POS tagging using deep learning.
pretrained Thai language model.
(SCB 10X, 2024) - open Thai LLM family.
(2023-2024) - open Thai instruction-tuned models.
open Southeast-Asian LMs covering Indonesian.
text corpus with 5M words and word segmentation.
dataset of speeches by the current Prime Minister of Thailand.
curated list of Ukrainian NLP datasets, models, etc.
curated list focused on machine translation and speech processing.
NLP library for Urdu.
POS, NER, and other NLP tasks.
Bilingual retrieval-evaluation benchmark for culturally grounded RAG. 400 rows, EN+UZ, MIT/CC-BY-4.0.
Vietnamese NLP toolkit.
Vietnamese text processing toolkit.
Vietnamese NLP toolkit.
Python Vietnamese core NLP toolkit.
on-device Vietnamese text-to-speech with voice cloning.
10K sentences for the constituency parsing task.
Vietnamese dependency treebank.
Vietnamese Universal Dependency Treebank.
free Vietnamese speech corpus, 15 hours of recorded speech (HCMUS AILab).
1.75M news sentences.
Vietnamese Text-to-SQL semantic parsing dataset (EMNLP-2020 Findings).
20M words across 15 bilingual books, 100 parallel English-Vietnamese texts, 250 parallel law texts, 5K news articles, and 2K film subtitles.
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…