跳到主要内容

multimodal

视觉语言、跨模态、图文对齐

17 篇,已拆 0 篇。回总览

2025

  • DAPO: An Open-Source LLM Reinforcement Learning System at Scale — paper=dapo@arXiv:2503.14476v2
  • SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training — paper=sft-memorizes-rl-generalizes-a-comparative@arXiv:2501.17161v2

2024

  • Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context — paper=gemini-1-5-unlocking-multimodal-understanding@arXiv:2403.05530v5

2023

  • BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models — paper=blip-2@arXiv:2301.12597v3
  • On the Importance of Noise Scheduling for Diffusion Models — paper=on-the-importance-of-noise-scheduling@arXiv:2301.10972v4
  • Segment Anything — paper=segment-anything@arXiv:2304.02643v1
  • Segment Anything Model for Medical Images? — paper=segment-anything-model-for-medical-images@arXiv:2304.14660v7
  • The Curse of Recursion: Training on Generated Data Makes Models Forget — paper=the-curse-of-recursion-training-on@arXiv:2305.17493v3

2022

  • Conditional Prompt Learning for Vision-Language Models — paper=conditional-prompt-learning-for-vision-language@arXiv:2203.05557v2

2021

  • Learning to Prompt for Vision-Language Models — paper=learning-to-prompt-for-vision-language@arXiv:2109.01134v6
  • Learning Transferable Visual Models From Natural Language Supervision — paper=clip@arXiv:2103.00020v1
  • NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction — paper=neus@arXiv:2106.10689v3

2020

  • An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — paper=vit@arXiv:2010.11929v2
  • Denoising Diffusion Probabilistic Models — paper=denoising-diffusion-probabilistic-models@arXiv:2006.11239v2

2019

  • Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model — paper=mastering-atari-go-chess-and-shogi@arXiv:1911.08265v2

2015

  • Deep Residual Learning for Image Recognition — paper=deep-residual-learning-for-image-recognition@arXiv:1512.03385v1
  • Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks — paper=unsupervised-representation-learning-with-deep-convolutional@arXiv:1511.06434v2