Skip to content

Entry

Attention Is All You Need

Appears in 6 awesome lists

(AIAYN) - Introducing multi-head self-attention neural networks with positional encoding to do sentence-level NLP without any RNN nor CNN - this paper is a must-read (also see this explanation and this visualization of the paper).

Open arxiv.org

Found in these lists

Awesome Artificial Intelligence

Section: Foundational papers · Introduced the Transformer architecture.

FreshScore 86

Awesome Deep Learning Resources

Section: Attention Mechanisms · (AIAYN) - Introducing multi-head self-attention neural networks with positional encoding to do sentence-level NLP without any RNN nor CNN - this paper is a must-read (also see this explanation and this visualization of the paper).

StaleScore 52

Awesome Forward Deployment Engineering (FDE)

Section: The Fundamental Whitepapers · The paper that started the Transformer/LLM revolution.

FreshScore 83

Awesome GPT Prompt Engineering

Section: Papers · Transformer introduction paper.

SlowScore 65

awesome-nlp

Section: Machine Translation · transformer; reset the field.

FreshScore 90

awesome-chatgpt

Section: The technical principle of ChatGPT · This paper introduces the structure of the original Transformer and is the basis for the Transformer family.

StaleScore 52

ReAct

The foundational paper defining the Thought/Action/Observation loop structure that underlies virtually every agent harness. Required reading for understanding why the loop is structured the way it is and where each harness component maps onto the reasoning-acting cycle.

In 7 listsDetails

Training language models to follow instructions with human feedback

This paper presents an RLHF approach to using supervised learning to fine-tuning. It is also known as a paper that illustrates the kernel of ChatGPT's thinking. Presumably, ChatGPT is an extended version of InstructGPT that enables fine-tuning on larger datasets.

In 6 listsDetails

Language Models are Few-Shot Learners

by Tom B. Brown (OpenAI) et al. - "We train GPT-3, an autoregressive language model with 175 billion parameters :scream:, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting."

In 4 lists

Direct Preference Optimization

Reframed preference alignment as a simple classification objective without explicit reward modelling.

In 3 lists

Constitutional AI

A method for training helpful and harmless AI assistants using written principles.

In 3 lists

Training Compute-Optimal Large Language Models

by Hoffmann et al. at DeepMind. TLDR: introduces a new 70B LM called "Chinchilla" that outperforms much bigger LMs (GPT-3, Gopher). DeepMind has found the secret to cheaply scale large language models — to be compute-optimal, model size and training data must be scaled equally. It shows that most…

In 3 lists

LoRA

and QLoRA - low-rank adapters and quantized fine-tuning; the standard for adapting LMs to NLP tasks on modest hardware.

In 2 lists

Retrieval-Augmented Generation

Combined parametric language models with external retrieval for knowledge-intensive tasks.

In 2 lists