PAPER EVIDENCE RANKING

论文证据排行榜

基于PaperMiner证据链动态评分。缺失证据保持未知,不等同于零分;不同任务论文不宜只凭总分直接比较。

综合侧重 科研价值 工程落地 综合影响
单项维度 可信度 可复现性 实验充分度 创新性 工程价值 影响力
评分版本:paper-evidence-v3 最低证据覆盖率:35% 至少3篇论文,取Top 20均值并校正样本量
排名 学校 得分 论文数 覆盖率
1 上海交通大学 DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation · Unsupervised Anomaly Detection in Brain MRI via Disentangled Anatomy Learning⋆ 86.19 274 100%
2 Tsinghua University ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving · OmniGAIA: Towards Native Omni-Modal AI Agents · LEGALONE: A FAMILY OF FOUNDATION MODELS FOR RELIABLE LEGAL REASONING 85.81 321 100%
3 Fudan University ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development · DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation 84.12 154 98%
4 Stanford University CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models · The limits of fair medical imaging AI in real-world generalization · What LLMs Think When You Don’t Tell Them What to Think About? 84.12 142 95%
5 Zhejiang University OmniGAIA: Towards Native Omni-Modal AI Agents · DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · On the Step Length Confounding in LLM Reasoning Data Selection 84.1 246 100%
6 The Chinese University of Hong Kong Approaching Low-Cost Cardiac Intelligence with Semi-Supervised Knowledge Distillation · Metamorphic Testing for Audio Content Moderation Software · StyleDoctor: Towards Specialist Reward Model for Style-centric Generation Tasks 83.68 145 96%
7 Massachusetts Institute of Technology The limits of fair medical imaging AI in real-world generalization · Path-RAG: Knowledge-Guided Key Region Retrieval for Open-ended Pathology Visual Question Answering · Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search 83.28 138 94%
8 University College INTRODUCTION AND NUMERICAL VALIDATION OF AN OPEN-SOURCE MATLAB PACKAGE FOR QUANTITATIVE ULTRASOUND TOMOGRAPHY VIA RAY-BORN INVERSION · Causal-Adversarial Probing of Clinical Covariates for Prostate MRI Grading · A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets 83.24 80 94%
9 Peking University DALD-PCAC: Density-Adaptive Learning Descriptor for Point Cloud Lossless Attribute Compression · MMAE: A Massive Multitask Audio Editing Benchmark · VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding 82.43 227 96%
10 National University of Singapore Certified Program Synthesis with a Multi-Modal Verifier · An Event-Based Opto-Tactile Skin · ALPaCA: Adapting Llama for Pathology Context Analysis to enable slide-level question answering 81.6 147 98%
11 Wuhan University Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation · C2|Q⟩: A Robust Framework for Bridging Classical and Quantum Software Development — RCR Report · Continual Vision-Language Learning for Remote Sensing: Benchmarking and Analysis 81.58 78 95%
12 University of Pennsylvania A Reproducible Evaluation of ANTs Similarity Metric Performance in Brain Image Registration · A Reproducible Evaluation of ANTs Similarity Metric Performance in Brain Image Registration · DECIPHERING SCIENTIFIC COLLABORATION IN BIOMEDICAL LLM RESEARCH: DYNAMICS, INSTITUTIONAL PARTICIPATION, AND RESOURCE DISPARITIES 81.57 52 93%
13 University of Michigan On the Step Length Confounding in LLM Reasoning Data Selection · Beyond Screenshots: Evaluating VLMs’ Understanding of UI Animations · EXP-Bench: Can AI Conduct AI Research Experiments? 81.17 75 93%
14 University of Chinese Academy of Sciences A Mutil-conditional Diffusion Transformer for Versatile Seismic Wave Generation · CL-VISTA: Benchmarking Continual Learning in Video Large Language Models · MGPC: Multimodal Network for Generalizable Point Cloud Completion With Modality Dropout and Progressive Decoding 81.15 139 93%
15 Sun Yat-sen University DepthAnything and SAM for UIE: Exploring Large Model Information Contributes to Underwater Image Restoration · DepthAnything and SAM for UIE: Exploring Large Model Information Contributes to Underwater Image Restoration · Comment Traps: How Defective Commented-out Code Augment Defects in AI-Assisted Code Generation 80.76 76 93%
16 Technical University of Munich The MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024: Efficient and Robust Aggregation Methods for Federated Learning · Coverage-Guided Road Selection and Prioritization for Efficient Testing in Autonomous Driving Systems · LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training 80.75 86 94%
17 Southeast University OmniGAIA: Towards Native Omni-Modal AI Agents · SpikACom: A Neuromorphic Computing Framework for Green Communications · V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views 80.74 67 94%
18 Imperial College SpikACom: A Neuromorphic Computing Framework for Green Communications · The MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024: Efficient and Robust Aggregation Methods for Federated Learning · InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries 80.73 85 92%
19 City University of Hong Kong Approaching Low-Cost Cardiac Intelligence with Semi-Supervised Knowledge Distillation · Metamorphic Testing for Audio Content Moderation Software · DALD-PCAC: Density-Adaptive Learning Descriptor for Point Cloud Lossless Attribute Compression 80.7 70 92%
20 Nanjing University DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding · RIVER: A REAL-TIME INTERACTION BENCHMARK FOR VIDEO LLMS 80.34 89 94%
21 Renmin University of China OmniGAIA: Towards Native Omni-Modal AI Agents · Metamorphic Testing for Audio Content Moderation Software · Feature Slice Matching for Precise Bug Detection 80.34 44 92%
22 Nanyang Technological University VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding · Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs · R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment 80.33 165 94%
23 University of Science and Technology of China DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · RIVER: A REAL-TIME INTERACTION BENCHMARK FOR VIDEO LLMS · NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results 80.32 183 96%
24 University of Cambridge SurvSurf: a partially monotonic neural network for first-hitting time prediction of intermittently observed discrete and continuous sequential events · ALPaCA: Adapting Llama for Pathology Context Analysis to enable slide-level question answering · PMPBench: A Paired Multi-Modal Pan-Cancer Benchmark for Medical Image Synthesis 80.23 88 94%
25 Beihang University ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving · LLMs-Powered Accurate Extraction, Querying and Intelligent Management of Literature-derived 2D Materials Data · NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results 79.92 84 92%
26 Johns Hopkins University Transfer Learning from One Cancer to Another via Deep Learning Domain Adaptation · RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology · SUMMARY OF THE INAUGURAL MUSIC SOURCE RESTORATION CHALLENGE 79.91 73 94%
27 Northwestern Polytechnical University A Generative Data Framework with Authentic Supervision for Underwater Image Restoration and Enhancement · Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset · OODEval: Evaluating Large Language Models on Object-Oriented Design 79.91 49 92%
28 University of California Path-RAG: Knowledge-Guided Key Region Retrieval for Open-ended Pathology Visual Question Answering · ConvRML: High-Quality Lensless Imaging with Random Multi-Focal Lenslets · Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games 79.9 223 94%
29 Carnegie Mellon University CONCUR: BENCHMARKING LLMS FOR CONCURRENT CODEGENERATION · HOTPOTQA: A Dataset for Diverse, Explainable Multi-hop Question Answering · OMTRA: A Multi-Task Generative Model for Structure-Based Drug Design 79.9 105 93%
30 The Hong Kong University of Science and Technology Metamorphic Testing for Audio Content Moderation Software · HiMed: Incentivizing Hindi Reasoning in Medical LLMs · Enhancing AIGC Service Efficiency with Adaptive Multi-Edge Collaboration in A Distributed System 79.49 162 94%
31 University of Electronic Science and Technology of China Decoding the Sequence Determinants of Locus-Specific DNA Methylation Across Human Tissues · Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis · DRNet: All-in-One Image Restoration via Prior-Guided Dynamic Reparameterization 79.47 67 93%
32 Beijing Institute of Technology Data Science and Technology Towards AGI Part I: Tiered Data Management · Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis · MM-DETR: An Efficient Multimodal Detection Transformer with Mamba-Driven Dual-Granularity Fusion and Frequency-Aware Modality Adapters 79.47 63 92%
33 Northeastern University From Noisy Historical Maps to Time-Series Oil Palm Mapping Without Annotation in Malaysia and Indonesia (2020–2024) · NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results · Scaling up fine-grained intracranial vessel annotations in computed tomography angiography 79.43 69 93%
34 University of Washington Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction · LandmarkLens: Predicting and Presenting Efective Landmarks for Mixed-Reality Urban Exploration · The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report 79.42 62 94%
35 The University of Hong Kong MatTools: Benchmarking Large Language Models for Materials Science Tools · LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback · PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis 79.07 80 91%
36 The Hong Kong Polytechnic University RESTORATION ADAPTATION FOR SEMANTIC SEGMENTATION ON LOW QUALITY IMAGES · NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results · High-resolution Photo Enhancement in Real-time: A Laplacian Pyramid Network 79.07 66 92%
37 University of Illinois Urbana-Champaign CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models · NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results · EXECUTIVE COUNTERFACTUALS: IMPROVING LLMS’ CAUSAL REASONING THROUGH CODE 79.07 65 91%
38 New York University Solaris: Building a Multiplayer Video World Model in Minecraft · Interactive Data Harmonization with LLM Agents: Opportunities and Challenges · Training Dynamics of Learning 3D-Rotational Equivariance 79.06 85 92%
39 Dalian University of Technology Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset · Revisiting Salient Object Detection from an Observer-Centric Perspective · The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report 79.06 42 92%
40 University of Oxford Connecting the Dots: A Machine Learning Ready Dataset for Ionospheric Forecasting Models · V-DPM: 4D Video Reconstruction with Dynamic Point Maps · FDIF: Formula-Driven Supervised Learning with Implicit Functions for 3D Medical Image Segmentation 79.03 84 93%
41 Tongji University ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving · An interactive enhanced driving dataset for autonomous driving · HiMed: Incentivizing Hindi Reasoning in Medical LLMs 78.8 61 93%
42 Beijing University of Posts and Telecommunications SIDME: Self-supervised Image Demoiréing via Masked EncoderDecoder Reconstruction · Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding · NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results 78.79 54 94%
43 Columbia University A Reproducible Evaluation of ANTs Similarity Metric Performance in Brain Image Registration · A Reproducible Evaluation of ANTs Similarity Metric Performance in Brain Image Registration · Visual Instruction Tuning 78.72 59 94%
44 Harbin Institute of Technology High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection · NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results · Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering 78.65 119 94%
45 Huazhong University of Science and Technology WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects · High-resolution Photo Enhancement in Real-time: A Laplacian Pyramid Network · Source-Free Domain Adaptation (SFDA) for Privacy-Preserving Seizure Subtype Classification 78.65 65 93%
46 Cornell University Tracking and Understanding Object Transformations · The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report · Beyond Failure Recovery: An Engagement-Aware Human-in-the-loop Framework for Robotic Systems 78.63 74 93%
47 Monash University High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection · Human Factors in Immersive Analytics · Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering 78.58 48 92%
48 Xi’an Jiaotong University Universally Unfiltered and Unseen: Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards · NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results · NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results 78.23 45 92%
49 University of Toronto LoRAFusion: Efficient LoRA Fine-Tuning for LLMs · Reasoning About Reasoning: Towards Informed and Reflective Use of LLM Reasoning in HCI · Velox: Learning Representations of 4D Geometry and Appearance 78.21 70 93%
50 Nanjing University of Science and Technology NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results · The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report · AIM 2025 challenge on Inverse Tone Mapping Report: Methods and Results 78.2 44 92%

排名用于发现值得进一步核查的论文,不替代同行评审。评分会随新增引用、复现、实验验证、新闻评价和落地证据持续更新。

社区反馈如何参与?

PaperMiner MCP已开放结构化论文反馈。可提交复现成功或失败、代码可运行或失效、Benchmark纠错、链接纠错和部署结果。

不直接改分反馈独立形成“社区信号”,与六维证据评分隔离。
必须带依据需要公开链接或实验结果哈希;无依据的点赞不进入系统。
自动交叉验证至少两个独立账号给出同类结论后,才标记为已有多人印证。
自动防刷个人密钥限频、去重;新账号和利益相关反馈自动降低权重。

MCP接口:submit_paper_feedbackget_paper_community_feedback

权限:有效PaperMiner账号获配的个人MCP密钥可提交;公共代理和共享只读密钥只能查询。每个账号每天最多提交5条新反馈。