pretraining
预训练、规模律、数据配方
120 篇,已拆 0 篇。回总览
2026
- DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection —
paper=deepseek-robustness-against-semantic-character-dual-space-mutated@arXiv:2604.12548v1 - DeepSeek-OCR 2: Visual Causal Flow —
paper=deepseek-ocr-2-visual-causal-flow@arXiv:2601.20552v1 - DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence —
paper=deepseek-v4-towards-highly-efficient-million-token-context@arXiv:2606.19348v1 - GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads —
paper=glm-5-serving-parameter-tuning-for-openclaw-single@arXiv:2607.02518v1 - GLM-5: from Vibe Coding to Agentic Engineering —
paper=glm-5-from-vibe-coding-to-agentic-engineering@arXiv:2602.15763v2 - GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents —
paper=glm-5v-turbo-toward-a-native-foundation-model@arXiv:2604.26752v3 - GLM-OCR Technical Report —
paper=glm-ocr-technical-report@arXiv:2603.10910v2 - GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation —
paper=glm-rag-graph-language-models-for-graph-based@arXiv:2607.28397v1 - Kimi K2.5: Visual Agentic Intelligence —
paper=kimi-k2-5-visual-agentic-intelligence@arXiv:2602.02276v2 - Kimi K3: Open Frontier Intelligence —
paper=kimi-k3-open-frontier-intelligence@arXiv:2607.24653v2 - Qwen Goes Brrr: Off-the-Shelf RAG for Ukrainian Multi-Domain Document Understanding —
paper=qwen-goes-brrr-off-the-shelf-rag-for@arXiv:2605.10296v1 - Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding —
paper=qwen-3d-a-generalist-3d-vision-language-model@arXiv:2608.02980v1 - Qwen-AgentWorld: Language World Models for General Agents —
paper=qwen-agentworld-language-world-models-for-general-agents@arXiv:2606.24597v1 - Qwen-Audio-3.0-Gen-Preview Technical Report —
paper=qwen-audio-3-0-gen-preview-technical-report@arXiv:2607.27011v2 - Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm —
paper=qwen-audio-3-0-tts-freely-controllable-and@arXiv:2607.23938v1 - Qwen-Audio-VAE Technical Report —
paper=qwen-audio-vae-technical-report@arXiv:2607.11738v1 - Qwen-BIM: developing large language model for BIM-based design with domain-specific benchmark and dataset —
paper=qwen-bim-developing-large-language-model-for-bim@arXiv:2602.20812v1 - Qwen-CUA: Native Computer Use for (almost) Everything —
paper=qwen-cua-native-computer-use-for-almost-everything@arXiv:2608.02352v1 - Qwen-Image-2.0 Technical Report —
paper=qwen-image-2-0-technical-report@arXiv:2605.10730v1 - Qwen-Image-2.0-RL Technical Report —
paper=qwen-image-2-0-rl-technical-report@arXiv:2606.27608v1 - Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation —
paper=qwen-image-agent-bridging-the-context-gap-in@arXiv:2606.26907v2 - Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation —
paper=qwen-image-bench-from-generation-to-creation-in@arXiv:2605.28091v2 - Qwen-Image-Flash: Beyond Objective Design —
paper=qwen-image-flash-beyond-objective-design@arXiv:2606.03746v2 - Qwen-Image-VAE-2.0 Technical Report —
paper=qwen-image-vae-2-0-technical-report@arXiv:2605.13565v1 - Qwen-Music Technical Report —
paper=qwen-music-technical-report@arXiv:2607.11699v2 - Qwen-MusicAVQA-7B: A Multimodal Model for Music Audio-Visual QA —
paper=qwen-musicavqa-7b-a-multimodal-model-for-music@arXiv:2608.11329v1 - Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models —
paper=qwen-robotmanip-technical-report-alignment-unlocks-scale-for@arXiv:2606.17846v2 - Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System —
paper=qwen-robotnav-technical-report-a-scalable-navigation-model@arXiv:2606.18112v3 - Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation —
paper=qwen-robotworld-technical-report-unifying-embodied-world-modeling@arXiv:2606.17030v3 - Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models —
paper=qwen-scope-turning-sparse-features-into-development-tools@arXiv:2605.11887v1 - Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents —
paper=qwen-ui-agent-technical-report-toward-next-generation@arXiv:2607.28227v1 - Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments —
paper=qwen-vla-unifying-vision-language-action-modeling-across@arXiv:2605.30280v2
2025
- DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models —
paper=deepseek-in-healthcare-a-survey-of-capabilities-risks@arXiv:2506.01257v1 - DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities —
paper=deepseek-on-a-trip-inducing-targeted-visual-hallucinations@arXiv:2502.07905v1 - DeepSeek performs better than other Large Language Models in Dental Cases —
paper=deepseek-performs-better-than-other-large-language-models@arXiv:2509.02036v1 - DeepSeek Powered Solid Dosage Formulation Design and Development —
paper=deepseek-powered-solid-dosage-formulation-design-and-development@arXiv:2503.11068v2 - DeepSeek reshaping healthcare in China's tertiary hospitals —
paper=deepseek-reshaping-healthcare-in-chinas-tertiary-hospitals@arXiv:2502.16732v2 - DeepSeek vs. ChatGPT vs. Claude: A Comparative Study for Scientific Computing and Scientific Machine Learning Tasks —
paper=deepseek-vs-chatgpt-vs-claude-a-comparative-study@arXiv:2502.17764v2 - DeepSeek-Inspired Exploration of RL-based LLMs and Synergy with Wireless Networks: A Survey —
paper=deepseek-inspired-exploration-of-rl-based-llms-and@arXiv:2503.09956v4 - DeepSeek-OCR: Contexts Optical Compression —
paper=deepseek-ocr-contexts-optical-compression@arXiv:2510.18234v1 - DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition —
paper=deepseek-prover-v2-advancing-formal-mathematical-reasoning-via@arXiv:2504.21801v2 - DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning —
paper=deepseek-r1-outperforms-gemini-2-0-pro-openai@arXiv:2502.17947v1 - DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning —
paper=deepseek-r1-thoughtology-lets-think-about-llm-reasoning@arXiv:2504.07128v3 - DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization? —
paper=deepseek-r1-vs-o3-mini-how-well-can@arXiv:2504.08120v3 - DeepSeek-V3, GPT-4, Phi-4, and LLaMA-3.3 generate correct code for LoRaWAN-related engineering tasks —
paper=deepseek-v3-gpt-4-phi-4-and-llama@arXiv:2502.14926v3 - DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models —
paper=deepseek-paradigm-shifts-and-technical-evolution-in-large@arXiv:2507.09955v1 - FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning —
paper=finescope-sae-guided-data-selection-enables@arXiv:2505.00624v3 - GLM Inference with AI-Generated Synthetic Data Using Misspecified Linear Regression —
paper=glm-inference-with-ai-generated-synthetic-data-using@arXiv:2503.21968v2 - GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models —
paper=glm-4-5-agentic-reasoning-and-coding-arc@arXiv:2508.06471v1 - GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning —
paper=glm-4-5v-and-glm-4-1v-thinking@arXiv:2507.01006v6 - GLM-TTS Technical Report —
paper=glm-tts-technical-report@arXiv:2512.14291v1 - Kimi k1.5: Scaling Reinforcement Learning with LLMs —
paper=kimi-k1-5-scaling-reinforcement-learning-with-llms@arXiv:2501.12599v4 - Kimi K2: Open Agentic Intelligence —
paper=kimi-k2-open-agentic-intelligence@arXiv:2507.20534v2 - Kimi Linear: An Expressive, Efficient Attention Architecture —
paper=kimi-linear-an-expressive-efficient-attention-architecture@arXiv:2510.26692v2 - Kimi-Audio Technical Report —
paper=kimi-audio-technical-report@arXiv:2504.18425v1 - Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents —
paper=kimi-dev-agentless-training-as-skill-prior-for@arXiv:2509.23045v3 - Kimi-VL Technical Report —
paper=kimi-vl-technical-report@arXiv:2504.07491v3 - Qwen it detect machine-generated text? —
paper=qwen-it-detect-machine-generated-text@arXiv:2501.09813v1 - Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information —
paper=qwen-look-again-guiding-vision-language-reasoning-models@arXiv:2505.23558v2 - Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems —
paper=qwen-vs-gemma-integration-with-whisper-a-comparative@arXiv:2506.13596v2 - Qwen-Image Technical Report —
paper=qwen-image-technical-report@arXiv:2508.02324v1 - Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition —
paper=qwen-image-layered-towards-inherent-editability-via-layer@arXiv:2512.15603v1 - Qwen3 Technical Report —
paper=qwen3-technical-report@arXiv:2505.09388v1 - Understanding R1-Zero-Like Training: A Critical Perspective —
paper=understanding-r1-zero-like-training-a@arXiv:2503.20783v2
2024
- AutoGLM: Autonomous Foundation Agents for GUIs —
paper=autoglm-autonomous-foundation-agents-for-guis@arXiv:2411.00820v1 - ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline —
paper=chatglm-math-improving-math-problem-solving-in-large@arXiv:2404.02893v1 - ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback —
paper=chatglm-rlhf-practices-of-aligning-large-language-models@arXiv:2404.00934v2 - ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools —
paper=chatglm-a-family-of-large-language-models-from@arXiv:2406.12793v2 - DeepSeek-V3 Technical Report —
paper=deepseek-v3-technical-report@arXiv:2412.19437v2 - DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models —
paper=deepseekmath-grpo@arXiv:2402.03300v3 - Docling Technical Report —
paper=docling-technical-report@arXiv:2408.09869v5 - From Local to Global: A Graph RAG Approach to Query-Focused Summarization —
paper=graphrag@arXiv:2404.16130v2 - GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot —
paper=glm-4-voice-towards-intelligent-and-human-like@arXiv:2412.02612v1 - Qwen2.5 Technical Report —
paper=qwen2-5-technical-report@arXiv:2412.15115v2
2023
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models —
paper=blip-2@arXiv:2301.12597v3 - CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X —
paper=codegeex-a-pre-trained-model-for-code-generation@arXiv:2303.17568v2 - CogVLM: Visual Expert for Pretrained Language Models —
paper=cogvlm-visual-expert-for-pretrained-language-models@arXiv:2311.03079v2 - Dense X Retrieval: What Retrieval Granularity Should We Use? —
paper=dense-x-retrieval-what-retrieval-granularity@arXiv:2312.06648v3 - GLM-Dialog: Noise-tolerant Pre-training for Knowledge-grounded Dialogue Generation —
paper=glm-dialog-noise-tolerant-pre-training-for-knowledge@arXiv:2302.14401v1 - GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints —
paper=gqa@arXiv:2305.13245v3 - Llama 2: Open Foundation and Fine-Tuned Chat Models —
paper=llama-2-open-foundation-and-fine@arXiv:2307.09288v2 - QLoRA: Efficient Finetuning of Quantized LLMs —
paper=qlora@arXiv:2305.14314v1 - Qwen Technical Report —
paper=qwen-technical-report@arXiv:2309.16609v1 - Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models —
paper=qwen-audio-advancing-universal-audio-understanding-via-unified@arXiv:2311.07919v2 - Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond —
paper=qwen-vl-a-versatile-vision-language-model-for@arXiv:2308.12966v3
2022
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure —
paper=bertopic@arXiv:2203.05794v1 - Conditional Prompt Learning for Vision-Language Models —
paper=conditional-prompt-learning-for-vision-language@arXiv:2203.05557v2 - Constitutional AI: Harmlessness from AI Feedback —
paper=constitutional-ai-harmlessness-from-ai-feedback@arXiv:2212.08073v1 - Efficient Few-Shot Learning Without Prompts —
paper=efficient-few-shot-learning-without-prompts@arXiv:2209.11055v1 - GLM for partially pooled categorical predictors with a case study in biosecurity —
paper=glm-for-partially-pooled-categorical-predictors-with-a@arXiv:2211.13848v1 - GLM-130B: An Open Bilingual Pre-trained Model —
paper=glm-130b-an-open-bilingual-pre-trained-model@arXiv:2210.02414v2 - Large Language Models are Zero-Shot Reasoners —
paper=zero-shot-cot@arXiv:2205.11916v4 - RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning —
paper=rlprompt@arXiv:2205.12548v3 - Self-Consistency Improves Chain of Thought Reasoning in Language Models —
paper=self-consistency@arXiv:2203.11171v4 - Training Compute-Optimal Large Language Models —
paper=chinchilla@arXiv:2203.15556v1
2021
- Generated Knowledge Prompting for Commonsense Reasoning —
paper=generated-knowledge-prompting-for-commonsense-reasoning@arXiv:2110.08387v3 - GLM: General Language Model Pretraining with Autoregressive Blank Infilling —
paper=glm@arXiv:2103.10360v2 - Learning to Prompt for Vision-Language Models —
paper=learning-to-prompt-for-vision-language@arXiv:2109.01134v6 - Learning Transferable Visual Models From Natural Language Supervision —
paper=clip@arXiv:2103.00020v1 - LoRA: Low-Rank Adaptation of Large Language Models —
paper=lora@arXiv:2106.09685v2 - Measuring Mathematical Problem Solving With the MATH Dataset —
paper=measuring-mathematical-problem-solving-with-the@arXiv:2103.03874v2 - Prefix-Tuning: Optimizing Continuous Prompts for Generation —
paper=prefix-tuning@arXiv:2101.00190v1 - Show Your Work: Scratchpads for Intermediate Computation with Language Models —
paper=show-your-work-scratchpads-for-intermediate@arXiv:2112.00114v1
2020
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale —
paper=vit@arXiv:2010.11929v2 - Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning —
paper=intrinsic-dimensionality-explains-the-effectiveness-of@arXiv:2012.13255v1 - Language Models are Few-Shot Learners —
paper=gpt-3@arXiv:2005.14165v4 - Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks —
paper=rag@arXiv:2005.11401v4 - Scaling Laws for Neural Language Models —
paper=scaling-laws-for-neural-language-models@arXiv:2001.08361v1
2019
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations —
paper=albert@arXiv:1909.11942v6 - Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer —
paper=exploring-the-limits-of-transfer-learning@arXiv:1910.10683v4 - HellaSwag: Can a Machine Really Finish Your Sentence? —
paper=hellaswag@arXiv:1905.07830v1 - Parameter-Efficient Transfer Learning for NLP —
paper=parameter-efficient-transfer-learning-for-nlp@arXiv:1902.00751v2 - RoBERTa: A Robustly Optimized BERT Pretraining Approach —
paper=roberta@arXiv:1907.11692v1 - Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks —
paper=sentence-bert@arXiv:1908.10084v1
2018
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding —
paper=bert@arXiv:1810.04805v2 - SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing —
paper=sentencepiece@arXiv:1808.06226v1 - Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates —
paper=subword-regularization-improving-neural-network-translation@arXiv:1804.10959v1 - Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge —
paper=think-you-have-solved-question-answering@arXiv:1803.05457v1
2016
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation —
paper=googles-neural-machine-translation-system-bridging@arXiv:1609.08144v2
2015
- Neural Machine Translation of Rare Words with Subword Units —
paper=neural-machine-translation-of-rare-words@arXiv:1508.07909v5