跳到主要内容

safety

对齐、越狱、隐私、幻觉与事实性

8 篇,已拆 0 篇。回总览

2026

  • How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms — paper=how-much-do-llms-hallucinate-in@arXiv:2603.08274v1

2024

  • Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone — paper=phi-3-technical-report-a-highly@arXiv:2404.14219v4

2023

  • Llama 2: Open Foundation and Fine-Tuned Chat Models — paper=llama-2-open-foundation-and-fine@arXiv:2307.09288v2
  • MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework — paper=metagpt@arXiv:2308.00352v7
  • Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — paper=not-what-youve-signed-up-for@arXiv:2302.12173v2
  • Ragas: Automated Evaluation of Retrieval Augmented Generation — paper=ragas@arXiv:2309.15217v2

2022

  • Training language models to follow instructions with human feedback — paper=instructgpt@arXiv:2203.02155v1

2015

  • Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks — paper=unsupervised-representation-learning-with-deep-convolutional@arXiv:1511.06434v2