Awesome AI Papers
Section: NLP
Entry
Appears in 6 awesome lists
This paper presents an RLHF approach to using supervised learning to fine-tuning. It is also known as a paper that illustrates the kernel of ChatGPT's thinking. Presumably, ChatGPT is an extended version of InstructGPT that enables fine-tuning on larger datasets.
Section: NLP
Section: Foundational papers · Established the instruction tuning and RLHF recipe used by InstructGPT.
Section: Papers · (InstructGPT)
Section: Instruction Tuning and Preference Optimization · training LMs to follow instructions with human feedback.
Section: Papers · by OpenAI. They call the resulting models InstructGPT. ChatGPT is a sibling model to InstructGPT.
Section: The technical principle of ChatGPT · This paper presents an RLHF approach to using supervised learning to fine-tuning. It is also known as a paper that illustrates the kernel of ChatGPT's thinking. Presumably, ChatGPT is an extended version of InstructGPT that enables fine-tuning on larger datasets.
Announcement of ChatGPT, a conversational model trained to answer follow-up questions, admit mistakes, challenge incorrect premises, and reject inappropriate requests. OpenAI blog, November 30, 2022.
The foundational paper defining the Thought/Action/Observation loop structure that underlies virtually every agent harness. Required reading for understanding why the loop is structured the way it is and where each harness component maps onto the reasoning-acting cycle.
(AIAYN) - Introducing multi-head self-attention neural networks with positional encoding to do sentence-level NLP without any RNN nor CNN - this paper is a must-read (also see this explanation and this visualization of the paper).
foundational result; intermediate reasoning steps improve performance.
by Tom B. Brown (OpenAI) et al. - "We train GPT-3, an autoregressive language model with 175 billion parameters :scream:, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting."
Reframed preference alignment as a simple classification objective without explicit reward modelling.
A method for training helpful and harmless AI assistants using written principles.
by Hoffmann et al. at DeepMind. TLDR: introduces a new 70B LM called "Chinchilla" that outperforms much bigger LMs (GPT-3, Gopher). DeepMind has found the secret to cheaply scale large language models — to be compute-optimal, model size and training data must be scaled equally. It shows that most…