efficiency
量化、蒸馏、剪枝、推理加速、长上下文
9 篇,已拆 0 篇。回总览
2026
- Models Take Notes at Prefill: KV Cache Can Be Editable and Composable —
paper=models-take-notes-at-prefill-kv@arXiv:2606.17107v1
2025
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models —
paper=deepseek-v3-2@arXiv:2512.02556v1 - FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning —
paper=finescope-sae-guided-data-selection-enables@arXiv:2505.00624v3
2024
- Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models —
paper=late-chunking-contextual-chunk-embeddings-using@arXiv:2409.04701v3
2023
- Lost in the Middle: How Language Models Use Long Contexts —
paper=lost-in-the-middle@arXiv:2307.03172v3 - MiniLLM: On-Policy Distillation of Large Language Models —
paper=minillm@arXiv:2306.08543v6
2022
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness —
paper=flashattention@arXiv:2205.14135v2
2018
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks —
paper=the-lottery-ticket-hypothesis-finding-sparse@arXiv:1803.03635v5
2015
- Distilling the Knowledge in a Neural Network —
paper=distilling-the-knowledge-in-a-neural@arXiv:1503.02531v1