PAPER EVIDENCE RANKING
基于PaperMiner证据链动态评分。缺失证据保持未知,不等同于零分;不同任务论文不宜只凭总分直接比较。
| 排名 | 学校 | 得分 | 论文数 | 覆盖率 |
|---|---|---|---|---|
| 1 | 上海交通大学 DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation · Unsupervised Anomaly Detection in Brain MRI via Disentangled Anatomy Learning⋆ | 86.19 | 274 | 100% |
| 2 | Tsinghua University ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving · OmniGAIA: Towards Native Omni-Modal AI Agents · LEGALONE: A FAMILY OF FOUNDATION MODELS FOR RELIABLE LEGAL REASONING | 85.81 | 321 | 100% |
| 3 | Fudan University ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development · DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation | 84.12 | 154 | 98% |
| 4 | Stanford University CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models · The limits of fair medical imaging AI in real-world generalization · What LLMs Think When You Don’t Tell Them What to Think About? | 84.12 | 142 | 95% |
| 5 | Zhejiang University OmniGAIA: Towards Native Omni-Modal AI Agents · DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · On the Step Length Confounding in LLM Reasoning Data Selection | 84.1 | 246 | 100% |
| 6 | The Chinese University of Hong Kong Approaching Low-Cost Cardiac Intelligence with Semi-Supervised Knowledge Distillation · Metamorphic Testing for Audio Content Moderation Software · StyleDoctor: Towards Specialist Reward Model for Style-centric Generation Tasks | 83.68 | 145 | 96% |
| 7 | Massachusetts Institute of Technology The limits of fair medical imaging AI in real-world generalization · Path-RAG: Knowledge-Guided Key Region Retrieval for Open-ended Pathology Visual Question Answering · Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search | 83.28 | 138 | 94% |
| 8 | University College INTRODUCTION AND NUMERICAL VALIDATION OF AN OPEN-SOURCE MATLAB PACKAGE FOR QUANTITATIVE ULTRASOUND TOMOGRAPHY VIA RAY-BORN INVERSION · Causal-Adversarial Probing of Clinical Covariates for Prostate MRI Grading · A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets | 83.24 | 80 | 94% |
| 9 | Peking University DALD-PCAC: Density-Adaptive Learning Descriptor for Point Cloud Lossless Attribute Compression · MMAE: A Massive Multitask Audio Editing Benchmark · VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding | 82.43 | 227 | 96% |
| 10 | National University of Singapore Certified Program Synthesis with a Multi-Modal Verifier · An Event-Based Opto-Tactile Skin · ALPaCA: Adapting Llama for Pathology Context Analysis to enable slide-level question answering | 81.6 | 147 | 98% |
| 11 | Wuhan University Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation · C2|Q⟩: A Robust Framework for Bridging Classical and Quantum Software Development — RCR Report · Continual Vision-Language Learning for Remote Sensing: Benchmarking and Analysis | 81.58 | 78 | 95% |
| 12 | University of Pennsylvania A Reproducible Evaluation of ANTs Similarity Metric Performance in Brain Image Registration · A Reproducible Evaluation of ANTs Similarity Metric Performance in Brain Image Registration · DECIPHERING SCIENTIFIC COLLABORATION IN BIOMEDICAL LLM RESEARCH: DYNAMICS, INSTITUTIONAL PARTICIPATION, AND RESOURCE DISPARITIES | 81.57 | 52 | 93% |
| 13 | University of Michigan On the Step Length Confounding in LLM Reasoning Data Selection · Beyond Screenshots: Evaluating VLMs’ Understanding of UI Animations · EXP-Bench: Can AI Conduct AI Research Experiments? | 81.17 | 75 | 93% |
| 14 | University of Chinese Academy of Sciences A Mutil-conditional Diffusion Transformer for Versatile Seismic Wave Generation · CL-VISTA: Benchmarking Continual Learning in Video Large Language Models · MGPC: Multimodal Network for Generalizable Point Cloud Completion With Modality Dropout and Progressive Decoding | 81.15 | 139 | 93% |
| 15 | Sun Yat-sen University DepthAnything and SAM for UIE: Exploring Large Model Information Contributes to Underwater Image Restoration · DepthAnything and SAM for UIE: Exploring Large Model Information Contributes to Underwater Image Restoration · Comment Traps: How Defective Commented-out Code Augment Defects in AI-Assisted Code Generation | 80.76 | 76 | 93% |
| 16 | Technical University of Munich The MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024: Efficient and Robust Aggregation Methods for Federated Learning · Coverage-Guided Road Selection and Prioritization for Efficient Testing in Autonomous Driving Systems · LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training | 80.75 | 86 | 94% |
| 17 | Southeast University OmniGAIA: Towards Native Omni-Modal AI Agents · SpikACom: A Neuromorphic Computing Framework for Green Communications · V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views | 80.74 | 67 | 94% |
| 18 | Imperial College SpikACom: A Neuromorphic Computing Framework for Green Communications · The MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024: Efficient and Robust Aggregation Methods for Federated Learning · InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries | 80.73 | 85 | 92% |
| 19 | City University of Hong Kong Approaching Low-Cost Cardiac Intelligence with Semi-Supervised Knowledge Distillation · Metamorphic Testing for Audio Content Moderation Software · DALD-PCAC: Density-Adaptive Learning Descriptor for Point Cloud Lossless Attribute Compression | 80.7 | 70 | 92% |
| 20 | Nanjing University DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding · RIVER: A REAL-TIME INTERACTION BENCHMARK FOR VIDEO LLMS | 80.34 | 89 | 94% |
| 21 | Renmin University of China OmniGAIA: Towards Native Omni-Modal AI Agents · Metamorphic Testing for Audio Content Moderation Software · Feature Slice Matching for Precise Bug Detection | 80.34 | 44 | 92% |
| 22 | Nanyang Technological University VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding · Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs · R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment | 80.33 | 165 | 94% |
| 23 | University of Science and Technology of China DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing · RIVER: A REAL-TIME INTERACTION BENCHMARK FOR VIDEO LLMS · NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results | 80.32 | 183 | 96% |
| 24 | University of Cambridge SurvSurf: a partially monotonic neural network for first-hitting time prediction of intermittently observed discrete and continuous sequential events · ALPaCA: Adapting Llama for Pathology Context Analysis to enable slide-level question answering · PMPBench: A Paired Multi-Modal Pan-Cancer Benchmark for Medical Image Synthesis | 80.23 | 88 | 94% |
| 25 | Beihang University ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving · LLMs-Powered Accurate Extraction, Querying and Intelligent Management of Literature-derived 2D Materials Data · NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results | 79.92 | 84 | 92% |
| 26 | Johns Hopkins University Transfer Learning from One Cancer to Another via Deep Learning Domain Adaptation · RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology · SUMMARY OF THE INAUGURAL MUSIC SOURCE RESTORATION CHALLENGE | 79.91 | 73 | 94% |
| 27 | Northwestern Polytechnical University A Generative Data Framework with Authentic Supervision for Underwater Image Restoration and Enhancement · Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset · OODEval: Evaluating Large Language Models on Object-Oriented Design | 79.91 | 49 | 92% |
| 28 | University of California Path-RAG: Knowledge-Guided Key Region Retrieval for Open-ended Pathology Visual Question Answering · ConvRML: High-Quality Lensless Imaging with Random Multi-Focal Lenslets · Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games | 79.9 | 223 | 94% |
| 29 | Carnegie Mellon University CONCUR: BENCHMARKING LLMS FOR CONCURRENT CODEGENERATION · HOTPOTQA: A Dataset for Diverse, Explainable Multi-hop Question Answering · OMTRA: A Multi-Task Generative Model for Structure-Based Drug Design | 79.9 | 105 | 93% |
| 30 | The Hong Kong University of Science and Technology Metamorphic Testing for Audio Content Moderation Software · HiMed: Incentivizing Hindi Reasoning in Medical LLMs · Enhancing AIGC Service Efficiency with Adaptive Multi-Edge Collaboration in A Distributed System | 79.49 | 162 | 94% |
| 31 | University of Electronic Science and Technology of China Decoding the Sequence Determinants of Locus-Specific DNA Methylation Across Human Tissues · Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis · DRNet: All-in-One Image Restoration via Prior-Guided Dynamic Reparameterization | 79.47 | 67 | 93% |
| 32 | Beijing Institute of Technology Data Science and Technology Towards AGI Part I: Tiered Data Management · Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis · MM-DETR: An Efficient Multimodal Detection Transformer with Mamba-Driven Dual-Granularity Fusion and Frequency-Aware Modality Adapters | 79.47 | 63 | 92% |
| 33 | Northeastern University From Noisy Historical Maps to Time-Series Oil Palm Mapping Without Annotation in Malaysia and Indonesia (2020–2024) · NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results · Scaling up fine-grained intracranial vessel annotations in computed tomography angiography | 79.43 | 69 | 93% |
| 34 | University of Washington Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction · LandmarkLens: Predicting and Presenting Efective Landmarks for Mixed-Reality Urban Exploration · The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report | 79.42 | 62 | 94% |
| 35 | The University of Hong Kong MatTools: Benchmarking Large Language Models for Materials Science Tools · LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback · PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis | 79.07 | 80 | 91% |
| 36 | The Hong Kong Polytechnic University RESTORATION ADAPTATION FOR SEMANTIC SEGMENTATION ON LOW QUALITY IMAGES · NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results · High-resolution Photo Enhancement in Real-time: A Laplacian Pyramid Network | 79.07 | 66 | 92% |
| 37 | University of Illinois Urbana-Champaign CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models · NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results · EXECUTIVE COUNTERFACTUALS: IMPROVING LLMS’ CAUSAL REASONING THROUGH CODE | 79.07 | 65 | 91% |
| 38 | New York University Solaris: Building a Multiplayer Video World Model in Minecraft · Interactive Data Harmonization with LLM Agents: Opportunities and Challenges · Training Dynamics of Learning 3D-Rotational Equivariance | 79.06 | 85 | 92% |
| 39 | Dalian University of Technology Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset · Revisiting Salient Object Detection from an Observer-Centric Perspective · The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report | 79.06 | 42 | 92% |
| 40 | University of Oxford Connecting the Dots: A Machine Learning Ready Dataset for Ionospheric Forecasting Models · V-DPM: 4D Video Reconstruction with Dynamic Point Maps · FDIF: Formula-Driven Supervised Learning with Implicit Functions for 3D Medical Image Segmentation | 79.03 | 84 | 93% |
| 41 | Tongji University ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving · An interactive enhanced driving dataset for autonomous driving · HiMed: Incentivizing Hindi Reasoning in Medical LLMs | 78.8 | 61 | 93% |
| 42 | Beijing University of Posts and Telecommunications SIDME: Self-supervised Image Demoiréing via Masked EncoderDecoder Reconstruction · Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding · NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results | 78.79 | 54 | 94% |
| 43 | Columbia University A Reproducible Evaluation of ANTs Similarity Metric Performance in Brain Image Registration · A Reproducible Evaluation of ANTs Similarity Metric Performance in Brain Image Registration · Visual Instruction Tuning | 78.72 | 59 | 94% |
| 44 | Harbin Institute of Technology High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection · NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results · Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering | 78.65 | 119 | 94% |
| 45 | Huazhong University of Science and Technology WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects · High-resolution Photo Enhancement in Real-time: A Laplacian Pyramid Network · Source-Free Domain Adaptation (SFDA) for Privacy-Preserving Seizure Subtype Classification | 78.65 | 65 | 93% |
| 46 | Cornell University Tracking and Understanding Object Transformations · The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report · Beyond Failure Recovery: An Engagement-Aware Human-in-the-loop Framework for Robotic Systems | 78.63 | 74 | 93% |
| 47 | Monash University High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection · Human Factors in Immersive Analytics · Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering | 78.58 | 48 | 92% |
| 48 | Xi’an Jiaotong University Universally Unfiltered and Unseen: Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards · NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results · NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results | 78.23 | 45 | 92% |
| 49 | University of Toronto LoRAFusion: Efficient LoRA Fine-Tuning for LLMs · Reasoning About Reasoning: Towards Informed and Reflective Use of LLM Reasoning in HCI · Velox: Learning Representations of 4D Geometry and Appearance | 78.21 | 70 | 93% |
| 50 | Nanjing University of Science and Technology NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results · The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report · AIM 2025 challenge on Inverse Tone Mapping Report: Methods and Results | 78.2 | 44 | 92% |
排名用于发现值得进一步核查的论文,不替代同行评审。评分会随新增引用、复现、实验验证、新闻评价和落地证据持续更新。
PaperMiner MCP已开放结构化论文反馈。可提交复现成功或失败、代码可运行或失效、Benchmark纠错、链接纠错和部署结果。
MCP接口:submit_paper_feedback、get_paper_community_feedback
权限:有效PaperMiner账号获配的个人MCP密钥可提交;公共代理和共享只读密钥只能查询。每个账号每天最多提交5条新反馈。