跳到主要内容

post-training

指令微调、RLHF、DPO 一族、推理模型

19 篇,已拆 0 篇。回总览

2026

  • Agent-Computer Observation Interfaces Enable Dynamic Computer Use — paper=agent-computer-observation-interfaces-enable-dynamic@arXiv:2606.29472v1
  • GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization — paper=gdpo@arXiv:2601.05242v1

2025

  • DAPO: An Open-Source LLM Reinforcement Learning System at Scale — paper=dapo@arXiv:2503.14476v2
  • DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — paper=deepseek-r1@arXiv:2501.12948v2
  • Group Sequence Policy Optimization — paper=group-sequence-policy-optimization@arXiv:2507.18071v2
  • MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention — paper=minimax-m1@arXiv:2506.13585v1
  • Rethinking Inference-Time Scaling: Efficiency Limits and Linguistic Signals — paper=rethinking-inference-time-scaling-efficiency-limits@arXiv:2504.14047v2
  • Understanding R1-Zero-Like Training: A Critical Perspective — paper=understanding-r1-zero-like-training-a@arXiv:2503.20783v2

2024

  • Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference — paper=chatbot-arena-an-open-platform-for@arXiv:2403.04132v1
  • ORPO: Monolithic Preference Optimization without Reference Model — paper=orpo@arXiv:2403.07691v2

2023

  • Direct Preference Optimization: Your Language Model is Secretly a Reward Model — paper=dpo@arXiv:2305.18290v3
  • Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — paper=judging-llm-as-a-judge-with@arXiv:2306.05685v4
  • Let's Verify Step by Step — paper=lets-verify-step-by-step@arXiv:2305.20050v1
  • Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — paper=not-what-youve-signed-up-for@arXiv:2302.12173v2

2022

  • Large Language Models Are Human-Level Prompt Engineers — paper=large-language-models-are-human-level@arXiv:2211.01910v2
  • Precise Zero-Shot Dense Retrieval without Relevance Labels — paper=hyde@arXiv:2212.10496v1
  • Training language models to follow instructions with human feedback — paper=instructgpt@arXiv:2203.02155v1

2020

  • Language Models are Few-Shot Learners — paper=gpt-3@arXiv:2005.14165v4

2017

  • Proximal Policy Optimization Algorithms — paper=proximal-policy-optimization-algorithms@arXiv:1707.06347v2