Skip to content
48

Awesome Question Answering

😎 A curated list of the Question Answering (QA)

769 stars104 forks124 entriesLast push Jan 13, 2022 (4 years ago)License CC0-1.0

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Recent Trends >Recent QA Models

paper:

github:

Demo:

paper:

github:

paper:

paper:

paper:

Recent Trends >Recent Language Models

ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

, Kevin Clark, et al., ICLR, 2020.

TinyBERT: Distilling BERT for Natural Language Understanding

, Xiaoqi Jiao, et al., ICLR, 2020.

MINILM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers

, Wenhui Wang, et al., arXiv, 2020.

T5: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

, Colin Raffel, et al., arXiv preprint, 2019.

In 3 lists

ERNIE: Enhanced Language Representation with Informative Entities

, Zhengyan Zhang, et al., ACL, 2019.

XLNet: Generalized Autoregressive Pretraining for Language Understanding

, Zhilin Yang, et al., arXiv preprint, 2019.

In 2 lists

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

, Zhenzhong Lan, et al., arXiv preprint, 2019.

RoBERTa: A Robustly Optimized BERT Pretraining Approach

, Yinhan Liu, et al., arXiv preprint, 2019.

In 2 lists

DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

, Victor sanh, et al., arXiv, 2019.

SpanBERT: Improving Pre-training by Representing and Predicting Spans

, Mandar Joshi, et al., TACL, 2019.

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

, Jacob Devlin, et al., NAACL 2019, 2018.

In 5 listsDetails

Recent Trends >AAAI 2020

paper:

Recent Trends >ACL 2019

Overview of the MEDIQA 2019 Shared Task on Textual Inference, Question Entailment and Question Answering

, Asma Ben Abacha, et al., ACL-W 2019, Aug 2019.

Towards Scalable and Reliable Capsule Networks for Challenging NLP Applications

, Wei Zhao, et al., ACL 2019, Jun 2019.

Cognitive Graph for Multi-Hop Reading Comprehension at Scale

, Ming Ding, et al., ACL 2019, Jun 2019.

Real-Time Open-Domain Question Answering with Dense-Sparse Phrase Index

, Minjoon Seo, et al., ACL 2019, Jun 2019.

Unsupervised Question Answering by Cloze Translation

, Patrick Lewis, et al., ACL 2019, Jun 2019.

SemEval-2019 Task 10: Math Question Answering

, Mark Hopkins, et al., ACL-W 2019, Jun 2019.

Improving Question Answering over Incomplete KBs with Knowledge-Aware Reader

, Wenhan Xiong, et al., ACL 2019, May 2019.

Matching Article Pairs with Graphical Decomposition and Convolutions

, Bang Liu, et al., ACL 2019, May 2019.

Episodic Memory Reader: Learning what to Remember for Question Answering from Streaming Data

, Moonsu Han, et al., ACL 2019, Mar 2019.

Natural Questions: a Benchmark for Question Answering Research

, Tom Kwiatkowski, et al., TACL 2019, Jan 2019.

Textbook Question Answering with Multi-modal Context Graph Understanding and Self-supervised Open-set Comprehension

, Daesik Kim, et al., ACL 2019, Nov 2018.

Recent Trends >EMNLP-IJCNLP 2019

Language Models as Knowledge Bases?

, Fabio Petron, et al., EMNLP-IJCNLP 2019, Sep 2019.

LXMERT: Learning Cross-Modality Encoder Representations from Transformers

, Hao Tan, et al., EMNLP-IJCNLP 2019, Dec 2019.

Answering Complex Open-domain Questions Through Iterative Query Generation

, Peng Qi, et al., EMNLP-IJCNLP 2019, Oct 2019.

KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning

, Bill Yuchen Lin, et al., EMNLP-IJCNLP 2019, Sep 2019.

Mixture Content Selection for Diverse Sequence Generation

, Jaemin Cho, et al., EMNLP-IJCNLP 2019, Sep 2019.

A Discrete Hard EM Approach for Weakly Supervised Question Answering

, Sewon Min, et al., EMNLP-IJCNLP, 2019, Sep 2019.

Recent Trends >Arxiv

Investigating the Successes and Failures of BERT for Passage Re-Ranking

, Harshith Padigela, et al., arXiv preprint, May 2019.

BERT with History Answer Embedding for Conversational Question Answering

, Chen Qu, et al., arXiv preprint, May 2019.

Understanding the Behaviors of BERT in Ranking

, Yifan Qiao, et al., arXiv preprint, Apr 2019.

BERT Post-Training for Review Reading Comprehension and Aspect-based Sentiment Analysis

, Hu Xu, et al., arXiv preprint, Apr 2019.

End-to-End Open-Domain Question Answering with BERTserini

, Wei Yang, et al., arXiv preprint, Feb 2019.

A BERT Baseline for the Natural Questions

, Chris Alberti, et al., arXiv preprint, Jan 2019.

Passage Re-ranking with BERT

, Rodrigo Nogueira, et al., arXiv preprint, Jan 2019.

SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering

, Chenguang Zhu, et al., arXiv, Dec 2018.

Recent Trends >Dataset

ELI5: Long Form Question Answering

, Angela Fan, et al., ACL 2019, Jul 2019

CODAH: An Adversarially-Authored Question Answering Dataset for Common Sense

, Michael Chen, et al., RepEval 2019, Jun 2019.

About QA >Analysis and Parsing for Pre-processing in QA systems

Morphological analysis

Systems

IBM Watson

Has state-of-the-arts performance.

Facebook DrQA

Applied to the SQuAD1.0 dataset. The SQuAD2.0 dataset has released. but DrQA is not tested yet.

MIT media lab's Knowledge graph

Is a freely-available semantic network, designed to help computers understand the meanings of words that people use.

In 2 lists

Publications

"Learning to Skim Text"

, Adams Wei Yu, Hongrae Lee, Quoc V. Le, 2017. : Show only what you want in Text

"Deep Joint Entity Disambiguation with Local Neural Attention"

, Octavian-Eugen Ganea and Thomas Hofmann, 2017.

"BI-DIRECTIONAL ATTENTION FLOW FOR MACHINE COMPREHENSION"

, Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, Hananneh Hajishirzi, ICLR, 2017.

"Capturing Semantic Similarity for Entity Linking with Convolutional Neural Networks"

, Matthew Francis-Landau, Greg Durrett and Dan Klei, NAACL-HLT 2016.

https://GitHub.com/matthewfl/nlp-entity-convnet

"Entity Linking with a Knowledge Base: Issues, Techniques, and Solutions"

, Wei Shen, Jianyong Wang, Jiawei Han, IEEE Transactions on Knowledge and Data Engineering(TKDE), 2014.

"Introduction to “This is Watson"

, IBM Journal of Research and Development, D. A. Ferrucci, 2012.

"A survey on question answering technology from an information retrieval perspective"

, Information Sciences, 2011.

"Question Answering in Restricted Domains: An Overview"

, Diego Mollá and José Luis Vicedo, Computational Linguistics, 2007

Codes

BiDAF

Bi-Directional Attention Flow (BIDAF) network is a multi-stage hierarchical process that represents the context at different levels of granularity and uses bi-directional attention flow mechanism to obtain a query-aware context representation without early summarization.; Official; Tensorflow v1.2

"BI-DIRECTIONAL ATTENTION FLOW FOR MACHINE COMPREHENSION"

, Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, Hananneh Hajishirzi, ICLR, 2017.

QANet

A Q&A architecture does not require recurrent networks: Its encoder consists exclusively of convolution and self-attention, where convolution models local interactions and self-attention models global interactions.; Google; Unofficial; Tensorflow v1.5; Paper

R-Net

An end-to-end neural networks model for reading comprehension style question answering, which aims to answer questions from a given passage.; MS; Unofficially by HKUST; Tensorflow v1.5

Paper

R-Net-in-Keras

R-NET re-implementation in Keras.; MS; Unofficial; Keras v2.0.6

DrQA

DrQA is a system for reading comprehension applied to open-domain question answering.; Facebook; Official; Pytorch v0.4; Paper

In 2 lists

BERT

A new language representation model which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations by jointly conditioning on both left and right context in all layers.;…

In 4 lists

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

, Jacob Devlin, et al., NAACL 2019, 2018.

In 5 listsDetails

Lectures

Question Answering - Natural Language Processing

By Dragomir Radev, Ph.D. | University of Michigan | 2016.

Slides

Question Answering with Knowledge Bases, Web and Beyond

By Scott Wen-tau Yih & Hao Ma | Microsoft Research | 2016.

Question Answering

By Dr. Mariana Neves | Hasso Plattner Institut | 2017.

Dataset Collections

NLIWOD's Question answering datasets

karthinkncode's Datasets for Natural Language Processing

Datasets

AI2 Science Questions v2.1(2017)

It consists of questions used in student assessments in the United States across elementary and middle school grade levels. Each question is 4-way multiple choice format and may or may not include a diagram element.

Paper:

Children's Book Test

CODAH Dataset

DeepMind Q&A Dataset; CNN/Daily Mail

Hermann et al. (2015) created two awesome datasets using news articles for Q&A research. Each dataset contains many documents (90k and 197k each), and each document companies on average 4 questions approximately. Each question is a sentence with one missing word/phrase which can be found from the…

In 2 lists

Paper:

ELI5

ELI5: Long Form Question Answering

, Angela Fan, et al., ACL 2019, Jul 2019

GraphQuestions

On generating Characteristic-rich Question sets for QA evaluation.

LC-QuAD

It is a gold standard KBQA (Question Answering over Knowledge Base) dataset containing 5000 Question and SPARQL queries. LC-QuAD uses DBpedia v04.16 as the target KB.

MS MARCO

This is for real-world question answering.

Paper:

MultiRC

A dataset of short paragraphs and multi-sentence questions

Paper:

NarrativeQA

It includes the list of documents with Wikipedia summaries, links to full stories, and questions and answers.

Paper:

NewsQA

A machine comprehension dataset

Paper:

Qestion-Answer Dataset by CMU

This is a corpus of Wikipedia articles, manually-generated factoid questions from them, and manually-generated answers to these questions, for use in academic research. These data were collected by Noah Smith, Michael Heilman, Rebecca Hwa, Shay Cohen, Kevin Gimpel, and many students at Carnegie…

SQuAD1.0

Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be…

In 4 listsDetails

Paper:

Paper:

Story cloze test

'Story Cloze Test' is a new commonsense reasoning framework for evaluating story understanding, story generation, and script learning. This test requires a system to choose the correct ending to a four-sentence story.

Paper:

TriviaQA

TriviaQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision…

In 2 lists

Paper:

WikiQA

A publicly available set of question and sentence pairs for open-domain question answering.

Datasets >The DeepQA Research Team in IBM Watson's publication within 5 years

"Unsupervised Entity-Relation Analysis in IBM Watson"

, Aditya Kalyanpur, J William Murdock, ACS, 2015.

"WatsonPaths: Scenario-based Question Answering and Inference over Unstructured Information"

, Adam Lally, Sugato Bachi, Michael A. Barborak, David W. Buchanan, Jennifer Chu-Carroll, David A. Ferrucci*, Michael R. Glass, Aditya Kalyanpur, Erik T. Mueller, J. William Murdock, Siddharth Patwardhan, John M. Prager, Christopher A. Welty, IBM Research Report RC25489, 2014.

"Medical Relation Extraction with Manifold Models"

, Chang Wang and James Fan, ACL, 2014.

Datasets >MS Research's publication within 5 years

"FigureQA: An Annotated Figure Dataset for Visual Reasoning"

, Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Akos Kadar, Adam Trischler, Yoshua Bengio, ICLR, 2018

"Stacked Attention Networks for Image Question Answering"

, Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, Alex Smola, CVPR, 2016.

"Question Answering with Knowledge Base, Web and Beyond"

, Yih, Scott Wen-tau and Ma, Hao, ACM SIGIR, 2016.

"NewsQA: A Machine Comprehension Dataset"

, Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, Kaheer Suleman, RepL4NLP, 2016.

"Table Cell Search for Question Answering"

, Sun, Huan and Ma, Hao and He, Xiaodong and Yih, Wen-tau and Su, Yu and Yan, Xifeng, WWW, 2016.

"WIKIQA: A Challenge Dataset for Open-Domain Question Answering"

, Yi Yang, Wen-tau Yih, and Christopher Meek, EMNLP, 2015.

"Web-based Question Answering: Revisiting AskMSR"

, Chen-Tse Tsai, Wen-tau Yih, and Christopher J.C. Burges, MSR-TR, 2015.

"Open Domain Question Answering via Semantic Enrichment"

, Huan Sun, Hao Ma, Wen-tau Yih, Chen-Tse Tsai, Jingjing Liu, and Ming-Wei Chang, WWW, 2015.

"An Overview of Microsoft Deep QA System on Stanford WebQuestions Benchmark"

, Zhenghao Wang, Shengquan Yan, Huaming Wang, and Xuedong Huang, MSR-TR, 2014.

Datasets >Google AI's publication within 5 years

"QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension"

, Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, Quoc V. Le, ICLR, 2018.

"Ask the Right Questions: Active Question Reformulation with Reinforcement Learning"

, Christian Buck and Jannis Bulian and Massimiliano Ciaramita and Wojciech Paweł Gajewski and Andrea Gesmundo and Neil Houlsby and Wei Wang, ICLR, 2018.

"Building Large Machine Reading-Comprehension Datasets using Paragraph Vectors"

, Radu Soricut, Nan Ding, 2018.

"An efficient framework for learning sentence representations"

, Lajanugen Logeswaran, Honglak Lee, ICLR, 2018.

"Did the model understand the question?"

, Pramod K. Mudrakarta and Ankur Taly and Mukund Sundararajan and Kedar Dhamdhere, ACL, 2018.

"Analyzing Language Learned by an Active Question Answering Agent"

, Christian Buck and Jannis Bulian and Massimiliano Ciaramita and Wojciech Gajewski and Andrea Gesmundo and Neil Houlsby and Wei Wang, NIPS, 2017.

"Learning Recurrent Span Representations for Extractive Question Answering"

, Kenton Lee and Shimi Salant and Tom Kwiatkowski and Ankur Parikh and Dipanjan Das and Jonathan Berant, ICLR, 2017.

"Neural Paraphrase Identification of Questions with Noisy Pretraining"

, Gaurav Singh Tomar and Thyago Duque and Oscar Täckström and Jakob Uszkoreit and Dipanjan Das, SCLeM, 2017.

Datasets >Facebook AI Research's publication within 5 years

Embodied Question Answering

, Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra, CVPR, 2018

Do explanations make VQA models more predictable to a human?

, Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh, EMNLP, 2018

Neural Compositional Denotational Semantics for Question Answering

, Nitish Gupta, Mike Lewis, EMNLP, 2018

Reading Wikipedia to Answer Open-Domain Questions

, Danqi Chen, Adam Fisch, Jason Weston & Antoine Bordes, ACL, 2017.

Building a Question-Answering System from Scratch— Part 1

Qeustion Answering with Tensorflow By Steven Hewitt, O'REILLY, 2017

Why question answering is hard

See category
94

Table of Contents

hesreallyhim/awesome-claude-code

A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…

Fresh★ 55k202 entriesPushed today
94

Awesome Agent Skills

VoltAgent/awesome-agent-skills

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

Fresh★ 35k839 entriesPushed today
93

Awesome Machine Learning

josephmisiti/awesome-machine-learning

A curated list of awesome Machine Learning frameworks, libraries and software.

Fresh★ 74k1188 entriesPushed 7 days ago
92

Awesome Production Machine Learning

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Fresh★ 21k519 entriesPushed 3 days ago
92

AWESOME DATA SCIENCE

academic/awesome-datascience

:memo: An awesome Data Science repository to learn and apply for real world problems.

Fresh★ 30k881 entriesPushed today
91

Static Analysis

analysis-tools-dev/static-analysis

⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…

Fresh★ 15k528 entriesPushed 8 days ago