跳到主要内容

efficiency

量化、蒸馏、剪枝、推理加速、长上下文

9 篇,已拆 0 篇。回总览

2026

  • Models Take Notes at Prefill: KV Cache Can Be Editable and Composable — paper=models-take-notes-at-prefill-kv@arXiv:2606.17107v1

2025

  • DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models — paper=deepseek-v3-2@arXiv:2512.02556v1
  • FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning — paper=finescope-sae-guided-data-selection-enables@arXiv:2505.00624v3

2024

  • Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models — paper=late-chunking-contextual-chunk-embeddings-using@arXiv:2409.04701v3

2023

  • Lost in the Middle: How Language Models Use Long Contexts — paper=lost-in-the-middle@arXiv:2307.03172v3
  • MiniLLM: On-Policy Distillation of Large Language Models — paper=minillm@arXiv:2306.08543v6

2022

  • FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — paper=flashattention@arXiv:2205.14135v2

2018

  • The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks — paper=the-lottery-ticket-hypothesis-finding-sparse@arXiv:1803.03635v5

2015

  • Distilling the Knowledge in a Neural Network — paper=distilling-the-knowledge-in-a-neural@arXiv:1503.02531v1