Update Dates 2609

2609 * *Computer Vision for Earth Observation Applications
* *Computer Vision for Geospatial Image Analysis
* *Computer Vision for Winter Sports
* *Event-based Vision in the Era of Generative AI - Transforming Perception and Visual Innovation Summary
* *FG
* *Foundational Models Beyond the Visible Spectrum
* *Generative AI for Photography
* *HARVEST-Vision: International Workshop on Applications of CV and HPC in Agriculture
* *Image/Video/Audio Quality in Computer Vision and Generative AI
* *International Workshop on Smart Waste Monitoring
* *Large Language and Vision Models for Autonomous Driving
* *LENS: Learning and Exploitation of Latent Space Geometries
* *Physical Retail AI
* *Pixels to Patients: Bridging CV State-of-Art with Clinical Impact
* *Real-World Surveillance: Applications and Challenges
* *Scene Graph for Structured Intelligencd
* *Synthetic & Adversarial ForEnsics
* *Synthetic Realities and Data in Biometric Analysis and Security
* *Video-based Human Recognition at Extreme Far Distances
* *Visual Art, Generative AI, and the Legal/Ethical Dilemma
* *WACV
* *WACV
* *Workshop on Computer Vision Systems for Document Analysis and Recognition
* *Workshop On Generative, Adversarial, Manipulation and Presentation Attacks In Biometrics
* *Workshop on Large Foundation Models in Biology and Biomedicine
* 1LoRa: Summation Compression for Very Low-Rank Adaptation
* 2COOOL: An Evaluation Benchmark for Generating Incident Reports on Out-of-Distribution Hazards in Autonomous Driving
* 2S-CEDiff: A Two-Stage Diffusion Framework for Generating High-Fidelity Contrast-Enhanced CT Images from Non-Contrast Scans
* 3-D Moiré Projection Model for Super-Resolution Point Localization
* 3D Cell Oversegmentation Correction via Geo-Wasserstein Divergence
* 3D Gaussian Point Encoders
* 3D Superquadric Splatting
* 3D Temporal Analysis for Autism Spectrum Disorder Screening During Attention Tasks
* 3DRealHead: Few-Shot Detailed Head Avatar
* 3DSceneEditor: Controllable 3D Scene Editing with Gaussian Splatting
* 4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
* 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
* 2D Gaussian Splatting, Two Dimensional Splatting (H4)
* Face Completion (H4)
* MRI Analysis for Brain, Cortex, Alzheimer's Disease (H4)
* Out of Distribution, OOD, Detection (H3)
* Seam Carving, Seam Carving Detection (H4)
* A-Predator: A Multibeam Echosounder Point Cloud Registration Network with Anisotropic Kernel Point Convolution
* A-V Representation Learning via Audio Shift Prediction for Multimodal Deepfake Detection and Temporal Localization
* ACBDT: SAR-Optical Cross-Modal Distillation for Sentinel-1/2 Building-Footprint Mapping in Heterogeneous Yangtze River Delta Cities
* Accelerated Dose Generation in Gamma Knife Radiosurgery Using a Wavelet Diffusion Model for Sparse Representation
* Achieving Text-Based Person Retrieval With Any Granularity
* Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
* ACuRE: Accurate Continuity-Regularized SpO2 Estimation Using Liquid Time-Constant Networks
* AD2: Analysis and Detection of Adversarial Threats in Visual Perception for End-to-End Autonomous Driving Systems
* Adapting Dense Vision-Language Relationships for Multi-Label Classification With Partial Label
* Adaptive Fusion of Multiple Land-Cover Products for Improved Spatial Representation of Key Land Classes in Central Asia
* Adaptive Hardness-Driven Dictionary Distillation for Incomplete Streaming View Clustering
* Adaptive Multi-Scale Feature Fusion for Paragraph-Level Handwritten Text Recognition via Spatial Attention
* Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
* Adaptive Segmented Doppler Compensation for Forward-Looking Radar Imaging
* Adaptive Sliding-Window Filtering for GNSS SPP-Aided Orbit Determination in Earth-Moon Space
* Adaptive Variational Inference: Beyond Bethe, Tree-Reweighted, and Convex Free Energies
* AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
* Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
* Advancing Player Identification and Tracking with Global ID Fusion (GIF)
* Advancing Precision Livestock Farming: Robust Country Chicken Detection via FeatherNet Fusion-YOLO and Hensense
* Advancing Robust Infrared Object Detection: Practical Insights on Model Development and Evaluation
* Adversarial Pseudo-replay for Exemplar-free Class-incremental Learning
* Adversarial-Robust Child Face Verification Using Spiking Neural Networks
* AEON: Adaptive Embedding Optimized Noise for Robust Watermarking in Diffusion Models
* Aerial Scene Classification Using Hierarchical and Multi-Stage Swin Transformer Features
* Aerosol Optical Depth Retrieval from MODIS Using a Physically Informed Machine Learning Framework
* AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-Training, Finetuning, and Evaluating Aerospace Embodied Foundation Models
* AFL-PRF: Adaptive Federated Learning for Low-Quality Data: Enhancing Performance, Robustness, and Fairness
* AFRAgent: An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
* AGCM: Attention-Guided Concept-based Model for Interpretable Affective Computing Applications
* AGENet: Adaptive Edge-aware Geodesic Distance Learning for Few-Shot Medical Image Segmentation
* Agentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging Systems
* AirLock+: Scaling UAV-to-Satellite Image Registration for Target Geolocalization and Geospatial Augmented Reality
* Algorithm for On-Sensor Agnostic Detection of Changes in Human Activity for Ultra-Low-Power Applications, An
* Align Video Diffusion Model with Online Video-Centric Preference Optimization
* Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
* Alignment and Distillation: A Robust Framework for Multimodal Domain Generalizable Human Action Recognition
* Altitude and Geographic Sensitivity Characteristics of the AIRS Satellite Spectrometer and Drift Correction Using Methane (CH4) Data
* AmbiGest: A Dataset of Social Gestures with Inter-Class Similarity and Intra-Class Variability
* Ambulance STARS: A Satellite-Driven Framework for Rapid Flood Impact Assessment and Time-Critical Ambulance Routing
* Analysis of Text Accuracy and Visual Alignment in Vision-Language Models for Artistic Text Generation
* Anatomically-guided masked autoencoder pre-training for aneurysm detection
* Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
* Anthropometrically Accurate Human Mesh Reconstruction via Body-Structured NeRF for Vision-Based Analysis
* Any Detector Can Detect Anything
* AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
* AnyBald: Toward Realistic Diffusion-Based Hair Removal In-The-Wild
* AortaDiff: A Unified Multitask Diffusion Framework for Contrast-Free AAA Imaging
* Applicability Assessment of Lutan-1 and Sentinel-1 for Potential Landslide Identification in Densely Vegetated Mountainous Areas: A Case Study of Hanyuan County, Sichuan Province, China
* Arc2Morph: Identity-Preserving Facial Morphing with Arc2Face
* ArchitectHead: Continuous Level of Detail Control for 3D Gaussian Head Avatars
* Are All Marine Species Created Equal? Performance Disparities in Underwater Object Detection
* ARGS: Advanced Regularization on Aligning Gaussians Over the Surface
* ARiSE: Efficient Mesh-Based Action Recognition from Wi-Fi Sensing on Edge Devices
* ART-ASyn: Anatomy-aware Realistic Texture-based Anomaly Synthesis Framework for Chest X-Rays
* ART: Actor-Related Tubelet for Detecting Complex-shaped Action Tubes
* ASC: Learning Augmentation Severity-Consistent Representations Improves Generalization via Augmentation Search
* Assessing Conservation Effectiveness of the Hainan Tropical Rainforest National Park Using Multi-Temporal Remote Sensing and Landscape Metrics
* Assessing the Potential of High-Resolution Multispectral and Structural Imagery for Plant Species Mapping in Mine Rehabilitation
* Assessing the Value of FY-4A/B Cloud-Top Height for Deep Learning-Based Tropical Cyclone Intensity Estimation over the Western North Pacific
* Assessing Vegetation-Hydrothermal Trend Regimes Across Elevation Gradients in Semiarid Mountains via Gaussian Mixture Models in Saudi Arabia
* Assimilation of FY-3G Precipitation Data Using a Machine Learning-Based Observation Operator in the CMA-MESO Regional Model
* ATFFormer: Asymmetric Temporal Fusion for Identity-Consistent Blind Video Face Restoration
* ATM: Enhanced Alignment for Text-to-Motion Generation
* Attention as Geometric Transformation: Revisiting Feature Distillation for Semantic Segmentation
* ATTN-FIQA: Interpretable Attention-based Face Image Quality Assessment with Vision Transformers
* Audio Deepfake Detectors vs. Real Fraud - the Fall of Benchmarks
* Auditable Knowledge-Graph Screening of Landslides for River Blockage and Dammed-Lake Assessment
* AugMapNet: Improving Spatial Latent Structure via BEV Grid Augmentation for Enhanced Vectorized Online HD Map Construction
* Augmenting with NeRFs: Fast Relocalization on Densified Datasets
* AusSmoke meets MultiNatSmoke: a fully-labelled diverse smoke segmentation dataset
* AuthGuard: Generalizable Deepfake Detection via Language Guidance
* Autocorrelation-based Fiducial Markers for Traceability
* Automated detection of positive and non-positive shyness in infants from videos
* Automated Pore Detection from In-Situ FDM 3D Printing Video: A Comparative Evaluation of Modern Segmentation Models
* Automated Suturing Skill Assessment in Robot-assisted Surgery from Endoscopic Videos using Clinically-guided Evaluation Criteria
* Automatic Identification and Assessment of Potential Geohazards in a Wide Area Based on Multisource Remote Sensing and Deep Learning
* Automating Tree Crown Delineation in UAV Orthomosaics Without Annotation: An Annotation-Free Framework Coupling DeepForest, Segment Anything, and Unsupervised Clustering
* Autoregressive Styled Text Image Generation, but Make it Reliable
* AutoSew: A Geometric Approach to Stitching Prediction with Graph Neural Networks
* AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization
* BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models
* BAFLE-DCT: Bypassing Adversarial Filters via Frequency-Selective Embedding in the DCT Domain
* Balancing the Scales: Uncertainty-Aware Two-Stage RL for Multiview Child Malnutrition Screening
* BanglaProtha: Evaluating Vision Language Models in Underrepresented Long-tail Cultural Contexts
* Basic: Bayesian Spiral Attention Classifier for Interpretable Medical Image Classification
* Bayesian Contrastive Augmented Open-set Action Recognition
* Bayesian Hierarchical Framework for Non-Linear and Iterative Transfer Learning of Gaussian Process, A
* Being Positive about Negative Queries: Exclusion Aware Multimodal Retrieval using Disentangled Representations
* Benchmarking Open-Access Building Footprints: A Multi-Dimensional Assessment with High-Fidelity References
* Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective
* Benchmarking Vision-Language Models for Traffic Scene Understanding in Inclement Winter Weather: The AWDB Benchmark
* Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
* Beyond Faces: A Multimodal Person Clustering for Unconstrained Environments
* Beyond Paired Data: Self-Supervised UAV Geo-Localization from Reference Imagery Alone
* Beyond Real Weights: Hypercomplex Representations for Stable Quantization
* Beyond Realism: Learning the Art of Expressive Composition with StickerNet
* Beyond Standalone Geo-Embeddings: Weighted Multi-Model Ensemble Prediction for Tropical Land-Cover Mapping
* Beyond the Encoder: Joint Encoder-Decoder Contrastive Pre-Training Improves Dense Prediction
* Beyond the Highlights: Video Retrieval with Salient and Surrounding Contexts
* Beyond the Surface: Incorporating 3D Facial Information into Action Unit Detection
* Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings
* Bi-Level Keypoint Relation Helps Versatile and Occluded Human Pose Estimation
* Bias-Aware Machine Learning Spatial Downscaling of GRACE Signals: Application to the Bug River Basin
* BiNAR: A Bi-Modal Framework for Non-Aligned RGB-IR 3D Reconstruction via Gaussian Splatting
* Biome-Specific Light Use Efficiency Model for Spatiotemporal Dynamics of GPP Across Europe Using PROBA-V and Sentinel-3 FAPAR, A
* BiPO: Bidirectional Partial Occlusion Network for Text-to-Motion Synthesis
* BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining
* BlendEmo: Structured Set Prediction and Pair-Conditioned Ratio Modeling for Blended Emotion Recognition
* Blur2Sharp: Human Novel Pose and View Synthesis with Generative Prior Refinement
* Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
* Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
* Boot-Shaped Terrain Screening and Deep Learning Semantic Segmentation for Landslide-Hazard Candidate Extraction from Airborne LiDAR DEM: A Case Study in Zhenxiong County, China
* BOP-Distrib: Revisiting 6D Pose Estimation Benchmarks for Better Evaluation under Visual Ambiguities
* Boundary-Sensitive Start-Time Estimation with Onset-centric Temporal Detection for Illegal Waste Dumping in Surveillance Video
* BoxSplitGen: A Generative Model for 3D Part Bounding Boxes in Varying Granularity
* BrandFusion: Aligning Image Generation with Brand Styles
* brat : Aligned Multi-View Embeddings for Brain MRI Analysis
* Breathing New Life Into Small Object Detection With Detection-Oriented Rectification
* BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries
* Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration
* Bridging Restoration and Diagnosis: A Comprehensive Benchmark for Retinal Fundus Enhancement
* Bridging the Domain Gap in Agricultural Vision: Parameter-Efficient VLM Adaptation via Expert Descriptions
* Bridging the Domain Gap in Small Multimodal Models: A Dual-level Alignment Perspective
* Bridging the Trust Gap: Interpretable AI for Clinical Decision Support in Arrhythmia Diagnostics
* BrightRate: Quality Assessment for User-Generated HDR Videos
* Broadcast2Pitch: Game State Reconstruction from Unconstrained Soccer Videos
* BrushEdit: All-in-One Image Inpainting and Editing
* Building Footprint Extraction in High-Density Urban Areas Based on Multi-Source Remote Sensing Data Fusion and ACM-PSPNet
* C2F-OR: Coarse-to-Fine Occlusion Removal for Facial Age Estimation on Partially Occluded Faces
* CAAC: Confidence-Aware Attention Calibration to Reduce Hallucinations in Large Vision-Language Models
* CADE: Continual Weakly-supervised Video Anomaly Detection with Ensembles
* CAFE: Cross-View Adaptive Fusion and Cluster Center Enhancement for Robust Multi-View Clustering
* CaFlow: Enhancing Long-Term Action Quality Assessment with Causal Counterfactual Flow
* CAKE: Context-Aware Kid Emotion In-the-Wild Dataset
* CalibBEV: LiDAR-Camera Calibration via BEV Alignment
* CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
* Can Image Splicing and Copy-Move Forgery Be Detected by the Same Model? Forensim: An Attention-Based State-Space Approach
* Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?
* Can You Find the Difference? Visually Identical Image Detection
* CANDOR: Centroid-Adjusted Normalized Decision-Level Optimization and Recognition for Blended Emotions
* CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
* CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
* CAPE: A Hierarchical Global-to-Local Framework for Crowd-Aware Pose Estimation
* CapStARE-LM: Capsule-based Spatiotemporal Architecture for Calibration-Free Gaze Estimation Using Facial Landmarks
* Capturing Individual Differences of Facial Expression for Authentic Expression Generation
* CARLA-Haze: A Synthetic Benchmark for Outdoor Image Dehazing
* CaRS: A Causal Intervention Segmentation Framework and Benchmark Dataset for Autonomous Driving under Transitional Weather Conditions
* CAST: Evaluating Multi-Object Trackers with Context-Aware Switch and Transfer Scores
* CasTex: Cascaded Text-to-Texture Synthesis via Explicit Texture Maps and Physically-Based Shading
* Causality-Driven Audits of Model Robustness
* CenterPoint-UAV: Context-Detail BEV Refinement for 3D Object Detection in UAV Point Clouds
* Centroiding Point-Objects With Event Cameras
* Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting
* ChameleonTuner: Automatic ISP Color Tuning in Subjective Scenarios
* Change Detection Network for Heterogeneous Remote Sensing Images Based on Decoupled Differential Architecture Search, A
* Change-Point-Based Deformation Grouping Strategy in Long-Term Near-Real-Time Deformation Monitoring, A
* ChartQA-X: Generating Explanations for Visual Chart Reasoning
* CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
* CL-LGFM: Early-Season Winter Wheat Mapping by Integrating Sentinel-2 NDVI and GPM Precipitation Data: A Case Study in the Chaohu Basin, China
* Clear Sights on Site: A Spatial-Adaptive Channel Network for Deblurring Construction Site Images
* CLIP's Visual Embedding Projector is a Few-shot Cornucopia
* CLIP-IT: CLIP-based Pairing of Histology Images with Privileged Textual Information
* CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering
* CLoCKDistill: Consistent Location and Context aware Knowledge Distillation for DETRs
* Closed-Form Statistical Expression for Evaluating Wind Speed and Direction Prediction Intervals from Doppler Lidar Arc Scans, A
* Cloud Occurrence, Phase, and Vertical Structure in a Dust-Influenced Eastern Mediterranean Region: Observations from Limassol, Cyprus
* CLUE: Bringing Machine Unlearning to Mobile Devices
* Cluster-Based Pseudo-Labeling for Semi-Supervised LiDAR Semantic Segmentation
* Cluster-Guided Adversarial Perturbations for Robust Contrastive Learning
* ClusterMine: Robust Label-Free Visual Out-Of-Distribution Detection via Concept Mining from Text Corpora
* Co-Burn: Combining dNBR Anchoring and Ordinal Learning for Cross-Event Fire Severity Mapping in New South Wales
* Co-STAR: Collaborative Curriculum Self-Training with Adaptive Regularization for Source-Free Video Domain Adaptation
* Coastal Vulnerability Index (CVI) Assessment of a Data-Sparse Delta: Quantifying the Contribution of InSAR-Derived Land Subsidence in the Volta Delta, Ghana
* CoDA: Color Distribution Probing for Efficient and Generalizable AI-Generated Image Detection
* Codebook Knowledge with Mamba-Transformer For Low-Light Image Enhancement
* CoL2A: Convolution-free Local Linear Attention for SpatioTemporal Event Processing
* Color Bind: Exploring Color Perception in Text-to-Image Models
* Color Preserving CMOS-SPAD Fusion for Multi-Frame HDR
* ColorPCR++: Unleash the Power of Color for Point Cloud Registration Using Hypergraph Computation
* CommonForms: A Large, Diverse Dataset for Form Field Detection
* Comp4D: Compositional 4D Scene Generation
* Comparing Modelled and Remotely Sensed Soil Moisture Products Using In Situ Observations in Liguria, Italy: Evaluation via SWI Filtering and Rescaling Techniques
* Complete Fuzzy Knowledge Representation With Knowledge Graph Embedding
* Complex-Valued HRU-Net with Cross-Gated Attention for PolSAR Semantic Segmentation
* Comprehensive Survey on Multimodal Recommender Systems: Taxonomy, Evaluation, and Future Directions, A
* Concord: Concept-Informed Diffusion for Dataset Distillation
* Conditional Text-to-Image Generation with Reference Guidance
* Confidence Through Parallel Attention for Depth and Uncertainty Estimation in Dynamic Environments
* Conjuring Positive Pairs for Efficient Unification of Representation Learning and Image Synthesis
* ConsensusXAI: A framework to examine class-wise agreement in medical imaging
* CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
* Context-Preserving Dermoscopic Editing: Mask-Guided Lesion-Aware Diffusion for Attribute Modification
* Continuous Satellite Monitoring of Reservoir Capacity Loss Using Deep Learning and Stochastic Mapping: The Poechos Reservoir and Regional Transferability in Northern Peru
* Contrast Then Confidence C^2: Contrastive Pretraining for Uncertainty-Aware Out-of-Distribution Detection in Satellite Imagery
* Contrastive Integrated Gradients: A Feature Attribution-Based Method for Explaining Whole Slide Image Classification
* ControlEvents: Controllable Synthesis of Event Camera Data with Foundational Prior from Image Diffusion Models
* Controllable Image Synthesis for Endoscopy: Leveraging Text and Spatial Guidance in Diffusion Models
* Controllable Long-term Motion Generation with Extended Joint Targets
* Controlled Accuracy Degradation of Photogrammetric 3D City Models
* Controlled Evaluation of Sentinel-2 Annual Compositing Strategies for Deep Learning-Based Mangrove Mapping in China
* Controlled Face Manipulation and Synthesis for Data Augmentation
* ControlVP: Interactive Geometric Refinement of AI-Generated Images with Consistent Vanishing Points
* Conversational Image Generation: Towards Multi-Round Personalized Generation with Multi-Modal Language Models
* CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
* CoreCaption: Core Caption based Text-to-Video Retrieval
* Correcting and Quantifying Systematic Errors in 3D Box Annotations for Autonomous Driving
* Cosine Similarity is Almost All You Need (for Prototypical-Part Models)
* Cost Savings from Automatic Quality Assessment of Generated Images
* Cost-Effective Approach to Estimate Quinoa Aboveground Biomass Volume Combining UAV RGB Data with Sentinel-1 and Sentinel-2 Satellite Imagery, A
* Cott-ADNet: Lightweight Real-Time Cotton Boll and Flower Detection Under Field Conditions
* Countering Multi-modal Representation Collapse through Rank-targeted Fusion
* CountingDINO A Training-free Pipeline for Class-Agnostic Counting using Unsupervised Backbones
* CP-VLM: Causal Prompting for Human Intention Inference with Vision-Language Models
* CPD-FCOS: A Scale-Isolated P2 Pathway for UAV Small-Object Detection
* Crafting Descriptive Information for a Zero-shot Method to Improve Knowledge-Based Visual Question Answering Performance
* Crafting Your Evolving Dreams: Concept-Incremental Versatile Customization
* CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion
* Crash2DocAI: Automated Integration of Post-Crash Car Part Images into Technical Reports
* CRISP: Cylindrical Rendering for In-Stream Point Clouds
* CropAT: Leveraging Diffusion-Generated Target-Like Cropped Objects for Pseudo-Label Refinement in Domain-Adaptive Object Detection
* Cross-Domain Mixup for Parcel-Level Crop Mapping on a Multi-Year Sentinel-2 Dataset from Slovakia
* Cross-Lingual Transfer for Complex Scripts: A Benchmark on End-to-End Khmer Scene Text Spotting
* Cross-Modal Event Encoder: Bridging Image-Text Knowledge to Event Streams
* Cross-Scale Performance Evaluation of GPM IMERG V07 Precipitation Products in a Typical Mountainous Monsoon Region
* CSF-Net: Context-Semantic Fusion Network for Large Mask Inpainting
* CSGaussian: Progressive Rate-Distortion Compression and Segmentation for 3D Gaussian Splatting
* CuDi: Curve Distillation for Efficient and Controllable Exposure Adjustment
* CURIO: Curvature-Aligned and Efficient OCR for Low-Resource Historical Manuscripts
* Curve Skeletonization in Continuous domain for Meshes and Point Clouds
* CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
* Cycle-Consistent Multi-Graph Matching for Self-Supervised Annotation of C. Elegans
* CycleSL: Server-Client Cyclical Update Driven Scalable Split Learning
* D2Mamba: Dual Domain Guided Informed Search in State Space Model for Underwater Image Enhancement
* DARB-Splatting: Generalizing Splatting with Decaying Anisotropic Radial Basis Functions
* Dark Noise Diffusion: Noise Synthesis for Low-Light Image Denoising
* Data-Driven Lipschitz Continuity: A Cost-Effective Approach to Improve Adversarial Robustness
* Data-Driven Loss Functions for Inference-Time Optimization in Text-to-Image
* Database-Agnostic Gait Enrollment using SetTransformers
* Dataset and Framework for Learning State-invariant Object Representations, A
* Dataset Pruning: Reducing Training Data by Examining SGD-Influence
* DATTA: Domain-Adversarial Test-Time Adaptation for Cross-Domain WiFi-Based Human Activity Recognition
* DBHN-Net: Dual-Branch Hybrid Neural Network for Low-Complexity Monaural Speech Enhancement
* DC4Former: Orientation-Stable UAV Disaster Image Segmentation via Diagonal-Complemented C4 Consistency
* dcFCI: Robust Causal Discovery Under Latent Confounding, Unfaithfulness, and Mixed Data
* DCSHARP: 3D Gaussian Splatting with Direction Cosine Spherical Harmonics and Shape-Aware Pruning
* DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
* DDEF-Net: A Difference-Guided Detail Enhancement Fusion Network for UAV-Based RGB-T Object Detection
* Deciphering Object Concepts: Hierarchical Cross-Modal Relational Reasoning for Mining Object-Attribute-Affordance Associations
* Decomposed Multi-Modality Fusion: Integrating Frames and Events for Efficient Visuomotor Policies
* Decomposition for Bayesian Networks: Local and Parallel Inference
* Decomposition Sampling for Efficient Region Annotations in Active Learning
* Decouple Then Converge: Handling Unknown Unlabeled Distributions in Long-Tailed Semi-Supervised Learning
* Decoupled Aquatic Greening and Water-Area Dynamics in Northeast Siberian Arctic Thermokarst Lakes from 2000 to 2025
* Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
* Decoupling Shape and Texture in SAM-2 via Controlled Texture Replacement
* Deep Image Decomposition for Medical Imaging Anonymization and Curation
* Deep Learning for Single-Frame Infrared Small and Dim Target Detection: A Paradigm-Oriented Review with Cross-Architecture Benchmarking
* Deep Learning for Water Body Segmentation in Remote Sensing Imagery: A Review
* Deep Learning-Based Monitoring of Tea Plant Growth and Nitrogen Status Using UAV Multisource Remote Sensing Features
* Deep Learning-Based Quantitative Precipitation Estimation Using Ground-Based Microwave Radiometer and Micro-Rain Radar Observations
* Deep Network for Object Detection on Inland Waters, A
* Deep Time Series Models: A Comprehensive Survey and Benchmark
* Deepfake Detection that Generalizes Across Benchmarks
* Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models
* Denoise, Divide, Distill, and Predict D3P: Towards Forecasting Long-horizon Real-world Anomaly from Normalcy
* DenseBEV: Transforming BEV Grid Cells into 3D Objects
* Depth Dynamics via One-Bit Frequency Probing in Embedded Direct Time-of-Flight Sensing
* Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection
* DermEVAL: A Dermatologist-Reviewed Benchmark for Multimodal Large Language Models
* Descrip3D: Enhancing Large Language Model-based 3D Scene Understanding with Object-Level Text Descriptions
* Detail Versus Certainty: A Metric Multi-Source Reconstruction of the Temple of Bel, Palmyra
* Detecting Deepfake Talking Heads from Facial Biometric Anomalies
* Detecting Object Tracking Failure via Sequential Hypothesis Testing
* Detecting Out-of-Distribution Objects through Class-Conditioned Inpainting
* Detecting Social Engagement of Elderly From Lifelog Image-streams to Identify Effective Cues for Autobiographic Recall
* Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
* Detector-Augmented SAMURAI for Long-Duration Drone Tracking
* Device-Robust Spectral Grading and Origin Detection from UV-VIS-NIR Images: Towards Practical Gemstone Quality Assessment
* DexAvatar: 3D Sign Language Reconstruction with Hand and Body Pose Priors
* DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
* DF-OOD: Real-Only Deepfake Detection via Confidence Dynamics under Perturbations
* DGCMam: Fusing Distance Graph Convolution and Mamba for Skeleton-based Action Recognition
* DGSRef: Decoupled Geometric-Semantic Refinement Network for High-Resolution Remote Sensing Segmentation
* Diagnose Like A REAL Pathologist: An Uncertainty-Focused Approach for Trustworthy Multi-Resolution Multiple Instance Learning
* DICE: Discrete Inversion Enabling Controllable Editing for Masked Generative Models
* DiffEyeSyn: User-specific Subtle Eye Movement Synthesis Using Diffusion Models
* DiffRegCD: Integrated Registration and Change Detection with Diffusion Features
* Diffuse4D: Completing Nerf-Stereo Depth via Diffusion-Driven Restoration in Dynamic Scenes
* Diffusion Noise Optimization for Synthetic VLM Training
* Diffusion-Based Action Recognition Generalizes to Untrained Domains
* Diffusion-Based Authentication of Copy Detection Patterns: A Multimodal Framework with Printer Signature Conditioning
* Digital Forensic AI You Can Explain: A Case Study on Video Source Camera Identification
* Digital Mapping of Soil and Water Indicators in Arid Regions Driven by High-Dimensional Environmental Covariates: A Comprehensive Evaluation of Metaheuristic Feature Selection and Hybrid Deep Learning Frameworks
* DiRe: Diversity-promoting Regularization for Dataset Condensation
* Direct Visual Grounding by Directing Attention of Visual Tokens
* DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment
* DiRIC: Diffusion Prior Refinement for Efficient Low-Rate Image Compression
* Discrepancy-Conditioned Residual Feature Refinement for Multi-Source Hyperspectral Classification
* Discrete Facial Encoding: A Framework for Data-driven Facial Display Discovery
* Disentangle and Regularize: Sign Language Production with Articulator-Based Disentanglement and Channel-Aware Regularization
* Disentangling Spectral and Environmental Controls on Inland River-Lake Water Quality Using Satellite Earth Observation-Driven Optimized Machine Learning
* DisenTS: Disentangled Channel Evolving Pattern Modeling for Multivariate Time Series Forecasting
* Distilling Diversity and Control in Diffusion Models
* Distilling Offline Action Detection Models into Real-Time Streaming Models
* Distilling What and Why: Enhancing Driver Intention Prediction with MLLMs
* Distribution Highlighted Reference-based Label Distribution Learning for Facial Age Estimation
* DiT-VTON: Diffusion Transformer Framework for Unified Multi-Category Virtual Try-On and Virtual Try-All with Integrated Image Editing
* Diurnal Asymmetry in the Relationships Between Urban Morphology and Canopy Urban Heat Islands: An Interpretable Machine Learning Analysis
* Diverse Sketch Colorization with Content-Enhanced Style Representation and Recolorization Distillation
* Diversity Preserving Coresets for Image Quality Assessment
* Divide and Refine: Enhancing Multimodal Representation and Explainability for Emotion Recognition in Conversation
* DM3Net: Dual-Camera Super-Resolution via Domain Modulation and Multi-scale Matching
* DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object Detection
* DMDNet: Decoupled Multimodal Detection Network for Fine-Grained Ulva Prolifera Segmentation
* DMS2F-HAD: A Dual-branch Mamba-based Spatial-Spectral Fusion Network for Hyperspectral Anomaly Detection
* DNA: Dual-branch Network with Adaptation for Open-Set Online Handwriting Generation
* Do generative video models understand physical principles?
* DocWaveDiff: A Predict-and-Refine approach for Document Image Enhancement with Wavelet U-Nets and Diffusion models
* DODA: Adapting Object Detectors to Dynamic Agricultural Environments in Real-Time with Diffusion
* Domain Generalizing DINO for Visual Regression via Latent Distractor Subspace Consistency
* DOODLE: Diffusion-based Out-of-Distribution Learning for Open-set LiDAR Semantic Segmentation
* DoTA: Latent Distribution Conditioned Data Attribution for Diffusion Models
* DOTGraph: CLIP-Driven Feature Disentanglement and Optimal Transport based Graph Learning for Few-Shot Segmentation
* DPBridge: Latent Diffusion Bridge for Dense Prediction
* Dragonite: Single-Step Drag-based Image Editing with Geometric-Semantic Guidance
* DREAM: Dynamic Prompts and GuidedMix for Efficient Continual Adaptation of Visual-Language Models
* DreamAnywhere: Object-Centric Panoramic 3D Scene Generation
* DreamCatcher: Efficient Multi-Concept Customization via Representation Finetuning
* DreamMakeup: Face Makeup Customization using Latent Diffusion Models
* Dressing the Imagination: A Dataset for AI-Powered Translation of Text into Fashion Outfits and A Novel NeRA Adapter for Enhanced Feature Adaptation
* DROMAL-Net: A Forest Larch Casebearer Detection Model Focusing on Global Feature Enhancement and Positional Dynamic Clustering
* Dronaquatics: Real-time Swimming Analytics Using Drone Captured Imagery
* DroneSplat+: Semantics-Enhanced 3D Gaussian Splatting for Robust 3D Reconstruction From In-the-Wild Drone Imagery
* DRWKV: Focusing on Object Edges for Low-Light Image Enhancement
* DSAN: Dual-Scale Aligned Network with Asymmetric Priors and Differentiable Soft-Edge Loss for SAR-to-Optical Image Translation
* DTMIR-Pro: Domain Translation with Prompt-based Latent-Space Generalization for Multi-Weather Image Restoration
* Dual-Domain Multimodal Hyperbolic Fusion for Cardiopulmonary Disease Diagnosis in Emergency Care
* DualGLEAN: Dual Allocation for VLM-Guided Generalized Category Discovery in Remote Sensing Images
* DualRes: Production-ready Dynamic Object Detection
* DUDA: Distilled Unsupervised Domain Adaptation for Lightweight Semantic Segmentation
* DuPLUS: Dual-Prompt Vision-Language Framework for Universal Medical Image Segmentation and Prognosis
* Dyna-Westdrive - VR-Based Multimodal Dataset for Emotion Recognition
* DynaGSLAM: Real-Time Gaussian-Splatting SLAM for Online Rendering, Tracking, Motion Predictions of Moving Objects in Dynamic Scenes
* Dynamic-Static Dataset of Facial Feature Variations Induced by Six Basic Tastes, A
* EA-VTON: Equivariance and AttentionFlow for Pose-Adaptive Virtual Try-On in Latent Diffusion Models
* EAR-Net: Pursuing End-to-End Absolute Rotations From Multi-View Images
* Earth-Limb-Constrained Framework for On-Orbit Geometric Calibration of GEO Wide-Field Area-Array Cameras, An
* Edge-Aware Image Manipulation via Diffusion Models with a Novel Structure-Preservation Loss
* Eff-GRot: Efficient and Generalizable Rotation Estimation with Transformers
* Efficient Deep Demosaicing with Spatially Downsampled Isotropic Networks
* Efficient Human Pose Estimation with Cross-Stage Fusion and Self-Distillation
* Efficient Inference for Large Reasoning Models: A Survey
* Efficient Isolated Sign Language Recognition via MediaPipe Landmarks and Affinity Mixed Attention: A Case Study On American Sign Language
* Efficient Multi-Rater Setup Towards Personalized and Diversified Medical Image Segmentation, An
* Efficient Text-Guided Convolutional Adapter for the Diffusion Model
* Efficient Vision Transformers via Token Merging with Head-Wise Attention Correction
* Efficient Visual Question Answering Pipeline for Autonomous Driving via Scene Region Compression
* Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models
* EGDNet: An Event-Guided Frequency-Aware Network with Cross-Modal Attention for Robust Weak Signature UAV Detection
* Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
* Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
* Egocentric Gesture Dataset for Robust Human-Robot Communication via Head Mounted Devices in Industrial and Military Settings
* EgoSSA: Egocentric Stereo Structure-Aware 3D Hand Reconstruction for American Sign Language Gesture Modeling
* Elastic Multi-Gradient Descent for Parallel Continual Learning
* Elastic Spiking Transformers for Efficient Gesture Understanding
* EllipssianNet: Image-guided Sampling of 2D Gaussians for Gaussian Splatting
* EMA: Effort Metric Attention for Anatomical Effort-Guided Human Motion Diffusion
* EmojiDiff: Advanced Facial Expression Control with High Identity Preservation in Portrait Generation
* Empirical Study of Siamese Vision Transformers for Scribe Re-Identification, An
* Employing Vision-Language Models for Face Image Quality Assessment
* Empowering Source-Free Domain Adaptation via MLLM-Guided Reliability-Based Curriculum Learning
* Enabling High-Quality In-the-Wild Imaging from Severely Aberrated Metalens Bursts
* ENCORE: A Neural Collapse Perspective on Out-of-Distribution Detection in Deep Neural Networks
* End-to-End Fine-Tuning of 3D Texture Generation using Differentiable Rewards
* EndoPBR: Photorealistic Synthetic Data for Surgical 3D Vision via Physically-based Rendering
* Enhanced 3D Lightning Localization for Low-Frequency Radio Observations over the Tibetan Plateau
* Enhanced Back-Projection of Vision Features for 3D Symmetry Detection
* Enhanced Nonlinear Grid Transformation Method for Weather Radar Echo Extrapolation, An
* Enhancement as Augmentation: Improving Detection in Highly Degraded Underwater Images Through Mixed-Domain Training
* Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
* Enhancing Object Detection Training via Joint Image-Annotation Generation
* Enhancing Reverse Distillation with Core Exemplar Learning for Unified Multi-Class Anomaly Detection
* Enhancing the Spatial Applicability of SMAP Soil Moisture Using Multi-Stage Machine Learning Downscaling
* Enhancing Vision Language Corruption Robustness using Cross-Distribution & Prompted Denoisers
* Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction
* Enhancing X-Ray Image Classification Through Heterogeneous Federated Learning With Natural Image-Augmented Models
* EPAT: Enhanced Parameter-Efficient Transfer Learning with Adaptive Strategy Fusion for Text-to-Image Person Re-identification
* Equivariant Sampling for Improving Diffusion Model-based Image Restoration
* eSkiTB: A Synthetic Event-Based Dataset for Tracking Skiers
* Estimating Annual Wildfire-Related Potential Above-Ground Biomass Loss in Eastern Canadian Boreal Forests Using Multi-Source Remote Sensing and XGBoost
* Estimating the Effects of Climate and Human Activities on Vegetation Coverage in Diverse Ecosystems: A Case Study of the Gansu Section of the Yellow River Basin
* Ethics-Aware Safe Reinforcement Learning for Rare-Event Risk Control in Interactive Urban Driving
* EUMD: Event-Based Unrolled Motion Deblurring Via Learnable Prior Optimization
* EUS-FAT: Video-Based Multiple Instance Learning for Liver Fat Quantification from Endoscopic Ultrasound
* EVA: Editing for Versatile Alignment Against Jailbreaks
* Evaluating Scale Transferability and Observation-Based Calibration in High-Resolution PM2.5 Downscaling
* Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
* Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
* Evaluating the Impact of DEM Resolution on Landslide Hazard Assessment: A Comparative Study Using LiDAR and 1:5000 Topographic Map-Derived DEMs
* Evaluation of the Fengyun-4B Downward Surface Shortwave Radiation (DSSR) Product over Guangxi Using a Dense Photovoltaic Station Network
* EVDI++: Event-Based Video Deblurring and Interpolation via Self-Supervised Learning
* Event-based Graph Representation with Spatial and Motion Vectors for Asynchronous Object Detection
* Event-based Liveness Detection using Temporal Ocular Dynamics: An Exploratory Approach
* Evolving Diverse Red-Team Language Models in Multi-Round Multi-Agent Games
* EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
* EX-FIQA: Leveraging Intermediate Early eXit Representations from Vision Transformers for Face Image Quality Assessment
* ExDDV: A New Dataset for Explainable Deepfake Detection in Video
* Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
* Exploiting Label-Independent Regularization from Spatial Patterns for Whole Slide Image Analysis
* Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
* Exploring Bayesian Prior-Driven Pseudo-Profile Reasoning for MLLM-based Micro-Expression Analysis
* Exploring Diffusion-generated Guidance for Thermal Image Super-resolution
* Exploring PPG-Guided Knowledge Distillation for Contactless Respiration Estimation
* Exploring the Boundaries of Diffusion Models for Offline Writer Identification with Sparse and Intra-Variable Data
* Exploring the Stochastic Regularisation in Normalisation Layers for Semi-Supervised Learning
* Exploring User Identification During Radar-Based Gesture Interaction
* Extreme Amodal Face Detection
* Eye-for-an-eye: Appearance Transfer with Dense Semantic Correspondence in Diffusion Models
* F-INR: Functional Tensor Decomposition for Implicit Neural Representations
* F-ViTA: Foundation Model Guided Visible-to-Infrared Translation
* Face Identity Unlearning for Retrieval via Embedding Dispersion
* Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
* FaceLayer: Multimodal 3D Face Asset Synthesis and Editing via Topology-Aware Disentangled Layer Composition
* Faces of Deception - a Multi-Level Analysis of Lip-Sync-Based Fake Ads Involving Politicians and Their Impact on Content Semantics and Recognition Systems
* Facial Expression Features of Deception in Dynamic Naturalistic Social Interactions
* FAE-Net: Fashion Attribute Editing via Disentangled Latent Conditioning in Diffusion Models
* FAIR-SIGHT: Fairness Assurance in Image Recognition via Simultaneous Conformal Thresholding and Dynamic Output Repair
* FairScene: Learning Class-Disentangled 2D/3D Representations for Semantic Scene Completion
* FairVLM: Enhancing Fairness and Prompt Sensitivity in Vision Language Models for Medical Image Segmentation
* FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs
* False Alarm Rectification for Early Smoke Segmentation
* FARF-Net: Frequency-guided Adaptive Receptive Field Network for Edge-enhanced Polyp Segmentation
* Fast 2DGS: Efficient Image Representation with Deep Gaussian Prior
* Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing
* Fast, Simple, and Flexible Scale Informative Feature Transform Module for Arbitrary Scale Image Super-Resolution, A
* FAST-EQA: Efficient Embodied Question Answering with Global and Local Region Relevancy
* FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
* FastPose-ViT: A Vision Transformer for Real-Time Spacecraft Pose Estimation
* FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature Banks
* FBR-DETR: An Efficient End-to-End Network for Real-Time Small-Object Detection in UAV Imagery
* FCC: Fully Connected Correlation for One-Shot Segmentation
* FD-MAD: Frequency-Domain Residual Analysis for Face Morphing Attack Detection
* FD-ProtoSCD: Semantic Change Detection in High-Resolution Remote Sensing Images via Frequency-Domain Disentanglement and Dynamic Prototype Learning
* Feature Alignment and Compositional Token for Human Pose Estimation
* Feature Inversion as a Lens on Vision Encoders
* Feature-Disentangling RGB-NIR Fusion Network for Remote Driver Physiological Measurement
* Feature-Space Planes Searcher: A Universal Domain Adaptation Framework for Interpretability and Computational Efficiency
* FedEFC: Federated Learning Using Enhanced Forward Correction Against Noisy Labels
* Federated Clustering: An Overview of Algorithm Evolution and Research Prospects
* Federated Learning via Variational Bayesian Inference: Personalization, Sparsity and Clustering
* Federated Model Synchronization for Diagnostic Redefinition through a Novel Selective Parameter Unlearning
* FedSCAl: Leveraging Server and Client Alignment for Unsupervised Federated Source-Free Domain Adaptation
* Feedback Alignment Meets Low-Rank Manifolds: A Structured Recipe for Local Learning
* Fetal and Neonatal Cortical Surface Reconstruction with Anatomical Normal-guidance and Perceptual Enhancements
* Few-Reference Identity-Aware Face Completion via Edge-Guided Structure and Self-Supervised Semantic Priors
* Few-Shot Shrub Identification via Hyperspectral Deep Learning: A Case Study of Caragana microphylla Lam in Shrub-Encroached Grasslands
* FFR-YOLO: A Frequency-Guided Fusion Reconstruction Network for Small-Object Detection in Remote Sensing Images
* FG-DINO: Inducing Foreground Approximation Into DINO for Low-Resolution Surveillance Videos
* FG-Tracer: Tracing Information Flow in Multimodal Large Language Models in Free-Form Generation
* FILD: Flash Interaction Latent Diffusion for Real-Time Text-Conditioned Human-Human Interaction Generation
* Fine-grained Defocus Blur Control for Generative Image Models
* FiNE: Fine-Grained Neuron-Level Model Editing for Reliable and Safe LLMs
* FLARES: Fast and Accurate LiDAR Multi-Range Semantic Segmentation
* FLoMo-Net: A Novel Task-Adaptive Mixture of Experts Routing Framework with Frequency and Uncertainty Correction for Medical Image Segmentation
* Flood-LDM: Generalizable Latent Diffusion Models for rapid and accurate zero-shot High-Resolution Flood Mapping
* Flow Augmentation and Knowledge Distillation for Lightweight Face Presentation Attack Detection
* FlowCLAS: Enhancing Normalizing Flow-Based Anomaly Segmentation Via Contrastive Learning
* FlowEO: Generative Unsupervised Domain Adaptation for Earth Observation
* FlowMorph: Revealing an Optimizable Flow Latent Space for Controlled Image Morphing
* FlyPose: Towards Robust Human Pose Estimation From Aerial Views
* FNOpt: Resolution-Agnostic, Self-Supervised Cloth Simulation using Meta-Optimization with Fourier Neural Operators
* FocalClick-XL: Toward Unified and High-Quality Interactive Segmentation
* FocalComm: Hard Instance-Aware Multi-Agent Perception
* Food Image Generation on Multi-Noun Categories
* Forensic Detection of Generated MRI Imagery Using Autoregressive Modeling and Frequency Analysis
* Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection
* ForestSplats: Deformable transient field for Gaussian Splatting in the Wild
* Forget Less by Learning Together through Concept Consolidation
* Foundation Models for Phenotyping Segmentation of Wheat Stripe Rust Resistance from UAV Hyperspectral Images
* Framework for Cross-Disaster Building Damage Assessment Using Cost-Sensitive Learning
* Framework for Real-Time Surgical Phase Recognition with Application to Robot-Assisted Partial Nephrectomy, A
* FreeCond: Free Lunch in the Input Conditions of Text-Guided Inpainting
* FreeKD+: A Frequency Knowledge Distillation Framework for Dense Prediction
* Frequency Is What You Need: Considering Word Frequency When Text Masking Benefits Vision-Language Model Pre-training
* Friend or Foe? Benchmarking Human Perception and ST-GCN Decoding of Embodied Social Intention
* FRLD-DF: Frequency Ring-Guided LoRA Adaptation of DINOv2 Vision Transformer for Generalizable Deepfake Detection
* From Bands to Depth: Understanding Bathymetry Decisions on Sentinel-2
* From Cognitive Priors to Instance Semantics: A Unified Framework for Multi-task Affective Computing
* From Darkness to Detail: Frequency-Aware SSMs for Low-Light Vision
* From Detection to Anticipation: Online Understanding of Struggles across Various Tasks and Activities
* From Downhill to Slalom: A Multi-Period Probabilistic Framework for Adaptive Snow Management in Alpine Ski Racing
* From Few-Shot to Zero-Shot Pallet Load Recognition: A Deployed Embedding-Based Vision System for Industrial Logistics
* From Filters to VLMs: Benchmarking Defogging Methods Through Object Detection and Segmentation Performance
* From Global to Granular: Revealing IQA Model Performance via Correlation Surface
* From Lightweight CNNs to SpikeNets: Benchmarking Accuracy-Energy Tradeoffs with Pruned Spiking SqueezeNet
* From Prompt to Production: Automating Brand-Safe Marketing Imagery with Text-to-Image Models
* From Real Faces to XR Avatars: Evaluating Face Recognition Vulnerability Through Avatar-Based Presentation Attacks
* From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
* From Storm Damage Detection to Windthrow Susceptibility Mapping: Evaluating Regional Transferability in Radiata Pine Plantations
* From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
* From Zero to Detail: A Progressive Spectral Decoupling Paradigm for UHD Image Restoration With New Benchmark
* FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder
* FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
* FujiView: Multimodal Late-Fusion for Predicting Scenic Visibility
* FuLLaMa: Training-free Diffusion-based Object Removal with Context Preservation
* Fully Unsupervised Self-debiasing of Text-to-Image Diffusion Models
* FUME: Fused Unified Multi-Gas Emission Network for Livestock Rumen Acidosis Detection
* Functional Differentiation of Postural Cues across Body Regions in Deception Detection: Implications for Human-Machine Trust and Interaction Systems
* FunFace: Feature Utility and Norm Estimation for Face Recognition
* FuseCLIP: Semantic-Guided Multidomain Fusion for Few-Shot Radar Active Jamming Recognition
* Fused Similarity Measure Based Alignment with Dual-Scale Adaptive Selection for Weakly Supervised Video Anomaly Detection
* GAEA: A Geolocation Aware Conversational Assistant
* GAITGen: Disentangled Motion-Pathology Impaired Gait Generative Model - Bringing Motion Generation to the Clinical Domain
* GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization
* Garberus: Three-Headed Video Classifier to Guard Against Illegal Waste Dumps
* GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
* Gated Temporal Fusion Transformers for Robust Multi-Object Tracking
* GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection
* GATEPose: A Graph Attention Transformer Enhanced with Pose and Orientation Angles for Pedestrian Crossing Intention Prediction
* Gaussian Adaptive Patching Powered Fully-Connected Spatial-Temporal Graph for Multivariate Time-Series Data
* Gaussian Representations for Video
* Gaussian Splatting Map Registration with Orthographic Bird's-Eye-View Renderings
* Gaussian Swaying?: Surface-Based Framework for Aerodynamic Simulation with 3D Gaussians
* GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting
* Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot Interaction
* GDoFS: Gaussian DoF Separation for Plausible 3D Geometry in Sparse-View 3DGS
* Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
* Gene-DML: Dual-Pathway Multi-Level Discrimination for Gene Expression Prediction from Histopathology Images
* GenEava: Generating Cartoon Avatars With Fine-Grained Facial Expressions From Realistic Diffusion-Based Faces
* General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
* Generalization of Real World Video Deblurring by Image-to-image Translation
* Generalized Category Discovery for LiDAR Semantic Segmentation
* Generalized Kullback-Leibler Divergence Loss
* Generalized Matérn Process for GNSS Coordinate Series Noise Modeling
* Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case Study
* Generative Data Augmentation for Skeleton Action Recognition
* Generic Deepfake Feature Space Discovery via Coarse-to-Fine Disentanglement Learning
* GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
* GenHSI: Controllable Generation of Human-Scene Interaction Videos
* Geo-DMAE: Geometric Deep Multi-Autoencoders for Monitoring Heterogeneous Normal Aging in Brain Subcortex
* Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
* GeoAI-Driven Wetland Change Analysis in the Sangamon River Watershed (2000-2025): A Comparative Assessment of Machine Learning and Deep Learning Approaches
* GeoHSAF: Geometric Hippocampus Shape Analysis Framework for Longitudinal Alzheimer's Disease Classification
* Geometry-Conditioned Diffusion for Occlusion-Robust In-Bed Pose Estimation
* Geometry-Constrained Framework for Automatic Geometric Positioning Accuracy Assessment of Large-Scale Satellite Imagery, A
* Geometry-Constrained Reference Sample Construction from Forest Inventory Compartments for Dominant Tree Species Mapping
* Geometry-Guided Semi-Supervised Multimodal Segmentation for UAV-Based Rice-Lodging Mapping
* Geospatial Foundation Models Improve Atoll Island Ecosystem Mapping: A Case Study Using AlphaEarth Embeddings
* GFE-Net: Geometry-Enhanced Feature Extraction Network for Semantic Segmentation of Large-Scale LiDAR Point Clouds
* GFT-GCN: Privacy-Preserving 3D Face Mesh Recognition with Spectral Diffusion
* GFT: Graph Feature Tuning for Efficient Point Cloud Analysis
* GHOST: Getting to the Bottom of Hallucinations with A Multi-round Consistency Benchmark
* GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model
* Global Attention-Based CNN for Interpretable Gender Classification in Palm Vein Biometrics
* Global Focal and Radial Distortion Averaging from Radial Fundamental Matrices for Robust Self-Calibration
* Global Navigation Satellite Systems (GNSS) in Climate Change Research: A Comprehensive Review
* GlossRefine: Gloss-Conditioned Transformer for Low-Resource Sign Language Motion Generation
* GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
* GPU-Based Solar Irradiance Estimation over Digital Surface Models Using Structurally Lossless Viewshed Compression
* GraDeCAR-Gradual Denoising by Contrastive Agreement-based Relabeling for Tackling Label Noise in Medical Imaging
* Gradient-Based Active Learning for Geospatial Semantic Segmentation with Large Vision Models
* Gradient-Free Classifier Guidance for Diffusion Model Sampling
* Gram-Schmidt Feature Reduction for Disentangling Style from Degradation in Historical Writer Identification
* GRAPE (Gaussian Rendering for Accelerated Pixel Enhancement) Brings Fast and Lightweight Arbitrary Super-Resolution
* Graph Query Networks for Object Detection with Automotive Radar
* Graph-Based Spectral Attention with Multi-Spectral Images for Illuminant Estimation
* GraspDiffusion: Synthesizing Realistic Whole-body Hand-Object Interaction
* GRAZE-FUSE: Grazing-Informed Integration of APSIM and Sentinel-2 for Pasture Biomass Monitoring
* GrounDiff: Diffusion-Based Ground Surface Generation from Digital Surface Models
* Grounding Degradations in Natural Language for All-In-One Video Restoration
* Grounding Descriptions in Images informs Zero-Shot Visual Recognition
* Groundwater Storage Dynamics and Attribution in the Wei River Basin Based on Dynamic Downscaling
* GroupPortrait: Multi-ID Portrait Generation with High Identity Preservation and Fine-Grained Control
* GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
* GSSP-KAN: An Efficient Kansformer-Based Network with Grouped Separable Sparse Convolution for Large-Scale LiDAR Point Cloud Semantic Segmentation
* GT-LandSDS: A Novel Spatiotemporal Integrated Framework for Land Use Simulation by Coupling Cellular Automata with Graph Attention Network and Transformer
* Guest Editors' Introduction: Special Section on Computational Photography (ICCP)
* Guided Model Merging for Hybrid Data Learning: Leveraging Centralized Data to Refine Decentralized Models
* Guided Texture Segmentation via Coordinate-Aware Class-Ratio Mapping
* Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
* GVHMR: Gravity-View Coordinates for Global Human Motion Recovery From Monocular Videos
* HABIT: Human Action Benchmark for Interactive Traffic in CARLA
* Handsurge: Localized Neural Surgery for Diffusion-Generated Hand Deformity Restoration
* HardFlow: Hard-Constrained Sampling for Flow-Matching Models via Trajectory Optimization
* Hardware-Aware Coding Function Design for Compressive Single-Photon 3D Cameras
* Harnessing Object Grounding for Time-Sensitive Video Understanding
* HDR Reconstruction Boosting with Training-Free and Exposure-Consistent Diffusion
* HDSMNet: Height-Guided Sparse Cross-Modal Fusion for High-Resolution Remote Sensing Semantic Segmentation
* HEART-PFL: Stable Personalized Federated Learning under Heterogeneity with Hierarchical Directional Alignment and Adversarial Knowledge Transfer
* Hestia: Voxel-Face-Aware Hierarchical Next-Best-View Acquisition for Efficient 3D Reconstruction
* Heteroscedastic Diffusion for Multi-Agent Trajectory Modeling
* Hierarchical Adaptive networks with Task vectors for Test-Time Adaptation
* Hierarchical Cross-Attention Transformer for Non-contact Multimodal Pain Classification using Remote Physiological Signals and Visual Features
* Hierarchical Fusion Method for SAR-Based Coastal Bathymetric Inversion in Short-Period Wave-Dominated Areas: A Case Study of the Wengtian Coast, Hainan Island
* Hierarchical Instance Tracking to Balance Privacy Preservation with Accessible Information
* HiFi-Deblur: High-Frequency Intense Image Deblurring with Frequency-Decoupled U-Net and Discrete Wavelet Transform
* High Fidelity and Real-time Video Face Swapping
* High-Level Semantics and Low-Level Features Fusion for Multi-Scale Object Detection in Dynamic Construction Environments
* High-Quality Entity Segmentation and Grounding
* High-Rate Mixout: Revisiting Mixout for Robust Domain Generalization
* High-Resolution 3D GPR Imaging of Concealed Surface Masonry in Pompeian Walls: Performance Analysis of Contact and Non-Contact Surveys
* High-Resolution Mapping of Forest Vegetation Types Using Multiplatform Imagery and Advanced Classification Techniques
* High-Spatiotemporal-Resolution Remote Sensing Retrieval of Evapotranspiration with Sentinel-2 Data by Sharpening MODIS Land Surface Temperature
* HiGlassRM: Learning to Remove High-prescription Glasses via Synthetic Dataset Generation
* HiMix: Hierarchical Visual-Textual Mixing Network for Lesion Segmentation
* Histogram Assisted Quality Aware Generative Model for Resolution Invariant NIR Image Colorization
* HistoMILKD: A Multiple Instance Learning based Multi-Teacher Knowledge Distillation Framework for Whole Slide Image Classification
* Histopath-C: Towards Realistic Domain Shifts for Histopathology Vision-Language Adaptation
* HNGT-Net: Hard-Negative Guided Topology Transfer for Lightweight Hyperspectral Small-Target Detection
* HodgeFormer: Transformers for Learnable Operators on Triangular Meshes through Data-Driven Hodge Matrices
* Holistic Method for Superquadric Fitting Using Unsupervised Clustering Analysis, A
* HOLO: Holistic Lightweight Optimization for Scene Understanding with Auto-Annotation and Multimodal Learning
* Homogeneous Terrain Unit Extraction by Integrating Superpixel Segmentation and Multiscale Region Merging: A Case Study in the Deeply Incised Valleys of Southeastern Tibet
* How Accurately Can Smartphone LiDAR Document the Exposed Coarse Root Architecture of Scots Pine? A Low-Cost Field Workflow
* How Far are We From Generating Missing Modalities With Foundation Models?
* How I Met Your Bias: Investigating Bias Amplification in Diffusion Models
* How to Design and Train Your Implicit Neural Representation for Video Compression
* HR2SIOD-CL: A Compressed Learning Framework for Object Detection in High-Resolution Remote Sensing Images
* HSAR-DETR: Hierarchical Spatial-Frequency Attention Network for UAV Small Object Detection
* HUI360: A 360° Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation
* Human knowledge integrated multi-modal learning for single source domain generalization
* Human Pose Aggregation for Multi-View Temporal Video Alignment
* HumanBench: Two Heads, No Legs, But Mostly Human, the State of Generative Capabilities in T2I Models
* HumanGuideNet: Adapter-Based Alignment of Deep Neural Networks with Human Similarity Judgments
* Hybrid Cloud Segmentation Approach Combining YOLOv8 Instance Segmentation with HSV Thresholding for Multi-Site Assessment
* Hybrid State Representation for Video Procedure Planning
* Hybrid Temporal-Spatial Novelty Detection for Illegal Waste Dumping in Surveillance Video
* HYMAVI: A Hybrid Mamba-Attention Network in Multi-View Framework for Volumetric Medical Image Segmentation
* HyPCA-Net: Advancing Multimodal Fusion in Medical Image Analysis
* HyperPose: Hyper-pose Embeddings for 3D-Aware Generative Models with Self-Supervised Disentangling of Pose and Scene
* Hyperspectral Technology: A Method Framework for the Estimation of Metal Content in Cobalt-Rich Crusts
* IB2MC: Information Bottleneck Inspired Balanced Multiview Clustering
* ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research
* IDEAL-M3D: Instance Diversity-Enriched Active Learning for Monocular 3D Detection
* Identification of Snowfall Riming and Aggregation Processes Using Ground-Based Triple-Frequency Radar
* Identity Verification from Human Scent using Channel Representation of 2D Gas Chromatography-Mass Spectrometry Data
* IDSync: Improving diffusion models through identity classification
* Ignoring the Decoy: Exposing and Tackling Forensic Distractions in Image Forgery Localization Using Masked Convolutions
* Illegal Waste Dumping Detection
* Illuminating Darkness: Learning to Enhance Low-light Images In-the-Wild
* Image Fusion Based on Prior Information
* Image Quality Assessment Methods for Multispectral Pan-Sharpening Images: A Comprehensive Review
* Image Restoration via Multi-Domain Learning
* Image-Guided Semantic Pseudo-LiDAR Point Generation for 3D Object Detection
* Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
* ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
* ImageNet-sES: A First Systematic Study of Sensor-Environment Simulation Anchored by Real Recaptures
* Imitating the Functionality of Image-to-Image Models Using a Single Example
* IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion
* iMotion-LLM: Instruction-Conditioned Trajectory Generation
* Impact of Radar-Constrained Effective Drop-Shape Relations on Polarimetric Radar Quantitative Precipitation Estimation in Typhoons
* IMPACT: Interpretable Most Important Person Analysis and Classification using Transformer-based Models
* improved architecture for part-based animal re-identification through semantic segmentation distillation, An
* Improved Land Surface Phenology Detection in China's Drylands and Associated Spatiotemporal Trends
* Improved Wildfire Spread Prediction with Time-Series Data and the WSTS+ Benchmark
* Improvement of Flood Risk Model Performance by Incorporating Sediment Factors
* Improving Animal Pose Estimation through Species Similarity Measures and Rigorous Label Definition
* Improving Backscatter-Based Surface Water Classification in Arid Environments Through Interferometric Coherence
* Improving Cross-River Turbidity Retrieval by Incorporating Environmental Variables: When and Why It Works
* Improving Negation Understanding in Medical Vision-Language Models via Contrastive Fine-Tuning
* Improving Out-of-Distribution Detection using Segmented Images and Cross-View Attention Fusion
* Improving Thermal Object Detection Robustness via Zero-Shot Inpainting-Based Data Augmentation
* Improving Viewpoint Robustness for Visual Recognition via Adversarial Training
* Improvise, Adapt, Overcome: Telescopic Adapters for Efficient Fine-tuning of Vision Language Models in Medical Imaging
* Inclusive AI for Group Interactions: Predicting Gaze-Direction Behaviors in People with Intellectual and Developmental Disabilities
* Industrial Brain: Self-Evolving Neuro-Symbolic Autonomy With Causal Resilience for Cyber-Physical Systems
* InfBA: Interference-Free Bottleneck Adaptation for Continual Learning
* Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
* Inpainting of Sparse Depth Maps from Monocular Depth-from-Focus on Pixel Processor Arrays
* INRetouch: Context Aware Implicit Neural Representation for Photography Retouching
* InstaFace: Identity-Preserving Facial Editing with Single Image Inference
* Instance-Dependent Early Stopping for Adaptive Data Pruning
* Instance-Level Cost-Sensitive Hypergraph Learning with Quality-Aware Structure Judgement for Anomaly Detection
* Instruction-Guided Distribution Maximization for General Few-Shot Intent Recognition
* Insulator-DETR: A Detection Transformer Tailored for Insulator Defect Inspection Based on UAV Remote Sensing
* Integrating Multi-scale and Multi-filtration Topological Features for Medical Image Classification*
* Inter-Stance: A Dyadic Multimodal Corpus for Conversational Stance Analysis
* InteracTalker: Prompt-Based Human-Object Interaction with Co-Speech Gesture Generation
* Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent Space
* Interactive Visual Exploration of Cross-Cultural Embodied Affect and Social Intent
* Interleaved Vision-and-Language Generation via Generative Voken
* Intra-Class Probabilistic Embeddings for Uncertainty Estimation in Vision-Language Models
* Intraoperative 2D/3D Registration via Spherical Similarity Learning and Differentiable Levenberg-Marquardt Optimization
* Investigating Bias and Fairness in Appearance-based Gaze Estimation
* Investigating the Relationship Between Micro-Expressions and Cognitive Load via a Novel Maze Paradigm
* iPathWS: An Integrated Scalable System for Explainable Pathology Diagnosis With Spectral Heterogeneity Engine Networks
* IPCD: Intrinsic Point-Cloud Decomposition
* IPTQ-ViT: Post-Training Quantization of Non-Linear Functions for Integer-Only Vision Transformers
* iSAGE: A Human-in-the-Loop Framework for Remote Sensing Semantic Segmentation via Sparse Point Supervision
* ISALux: Illumination and Semantics-Aware Transformer Employing Mixture of Experts for Low Light Image Enhancement
* Isolating the Role of Temporal Information in Video Saliency: A Controlled Experimental Analysis
* ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
* Jacobians of Light Scattering Properties Based on the Improved Geometric Optics Method (IGOM)
* JetBench: Quality-Aware Benchmarking of Vision Models for Jet Parameter Classification in Heavy-Ion Physics
* JOCA: Task-Driven Joint Optimisation of Camera Hardware and Adaptive Camera Control Algorithms
* Joint Modeling of Corruption-Driven and Information-Limited Uncertainty for Robust 3D Gaussian Splatting
* Joint Optimization of Camera Model and Deep Neural Network for Image Recognition
* K-Fold Ensemble of Hiera and VideoMAE for Illegal Waste Dumping Detection
* K-Track: Kalman-Enhanced Tracking for Accelerating Deep Point Trackers on Edge Devices
* K-Vehicles: A Remote Sensing Dataset for Vehicle Detection in Aerial Imagery
* KD360-VoxelBEV: LiDAR and 360-degree Camera Cross Modality Knowledge Distillation for Bird's-Eye-View Segmentation
* Kernel PCA for Out-of-Distribution Detection: Non-Linear Kernel Selection and Approximation
* KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
* King(dom) is Naked: Lightweight Machine Learning for Hyperspectral Bare Soil Detection, The
* KMOPS: Keypoint-Driven Method for Multi-Object Pose and Metric Size Estimation from Stereo Images
* Knowledge Diffusion-Based Adaptive Alignment With Hierarchical Context for Video Temporal Grounding
* Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
* Land-Cover and Land-Use Mapping Under Limited Data Highlights Hyperparameter Stability and Predictor Design
* Landscape Greening Following Unseasonal Precipitation Along a Desert-Alpine Gradient
* LangPose: Language-Aligned Motion for Robust 3D Human Pose Estimation
* LangSurf: Language-Embedded Surface Gaussians for 3D Scene Understanding
* Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression
* Language-4D Cross-Boosting for Generalized Zero-Shot 6DoF Tracking and 3D Reconstruction
* Large Sign Language Models: Toward 3D American Sign Language Translation
* Large-Scale 3D Representation Dataset and Benchmark for Continuous Sign Language Understanding, A
* LASER: Lip Landmark Assisted Speaker Detection for Robustness
* LASOR: Towards Clinically Transparent and Explainable Ophthalmic Report Generation via Lesion-Aware Segmentation
* Latent Space Manifold Geometry for Robust Image Classification
* Latent Uncertainty-Aware Multi-View SDF Scan Completion
* LaVIDE: Language-Prompted Satellite Change Detection via Map-Image Alignment
* Layout Anything: One Transformer for Universal Room Layout Estimation
* Layout Optimization of Urban Emergency Shelter Sites Under Compound Disaster Scenarios Based on MOGWO
* LBOR: Laplace-Beltrami Operator Regularization for Robust Skeleton-based Isolated Sign Language Recognition
* Learnable Query-Enhanced Pose Transformation
* Learned Off-Aperture Encoding for Wide Field-of-View RGBD Imaging
* Learning Action Hierarchies via Hybrid Geometric Diffusion
* Learning Beyond Labels: Self-Supervised Handwritten Text Recognition
* Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
* Learning Domain Agnostic Latent Embeddings of 3D Faces for Zero-Shot Animal Expression Transfer
* Learning From Crowds With Multiple Feature Dynamic Fusion-Based Annotation Generation
* Learning from Unknown for Open-Set Test-Time Adaptation
* Learning Group Actions In Disentangled Latent Image Representations
* Learning Mask-Aware Offsets: Two-branch Deformable Attention Networks for Inpainting with Masked Region Avoidance
* Learning Shape Anchors for Holistic Indoor Scene Understanding
* Learning spatio-temporal feature representations for video-based gaze estimation
* Learning Subglacial Bed Topography from Sparse Radar with Physics-Guided Residuals
* Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
* Learning Topology-Aware Representations via Test-Time Adaptation for Anomaly Segmentation
* Learning Unified Spatio-temporal Representations for Efficient Compressed Video Understanding
* Learning When to Adapt: Forecast-Driven Meta Learning for Few-Shot Professional Action Recognition under Data Scarcity
* LENVIZ: A High-Resolution Low-Exposure Night Vision Benchmark Dataset
* Less is More: Agentic Prompt Design for Safe VLM Action Selection
* Leveraging Pretrained Representations for Cross-Modal Point Cloud Completion
* Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion Models
* Leveraging Sparsity for Privacy in Collaborative Inference
* Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking
* LibreFace 2.0: A Generalizable Facial Expression Analysis Toolkit Leveraging Synthetic Data
* LiDAR-DHMT: LiDAR-Adaptive Dual Hierarchical Mask Transformer for Robust Freespace Detection and Semantic Segmentation
* LIFT+: Lightweight Fine-Tuning for Long-Tail Learning
* LightGazeNet: A Lightweight GNN-based Architecture for Gaze Estimation
* LighthouseGS: Indoor Structure-aware 3D Gaussian Splatting for Panorama-Style Mobile Captures
* Lightweight Multi-Scale Fusion for Real-Time Autonomous Driving Segmentation
* Lightweight Near-Infrared Spectral Reconstruction from Red UAV Imagery Using Artificial Intelligence for Low-Cost Remote Sensing
* Lightweight Temporal Detection Framework for Illegal Waste Dumping in Real Surveillance Footage, A
* Line Art Colorization with Offset Prior-based Diffusion Model
* Linear Attention Framework with Dual-Axis Multi-Scale Fusion for Fine-Grained Eucalyptus Change Detection, A
* Linearly Solving Robust Rotation Estimation
* LiteViT-CAM: A Lightweight and Intrinsically Explainable Framework for Realtime Clinical Video Analysis on Edge Devices
* Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback, A
* LLM Augmented Intervenable Multimodal Adaptor for Post-operative Complication Prediction in Lung Cancer Surgery
* Locally Explaining Prediction Behavior via Gradual Interventions and Measuring Property Gradients
* LogicCBMs: Logic-Enhanced Concept-Based Learning
* Logit-Adjusted Test-Time Adaptation under Partial Class Imbalance
* Long&short Exposures Guided Diffusion Model for Realistic Local Motion Deblurring
* Long-Term Dynamics of Land Degradation Risk in Arid Northwest China Revealed by an Integrated Risk Index
* Long-Term InSAR Monitoring and Anomaly Detection of Railway Deformation in Shanghai
* LooC: Effective Low-Dimensional Codebook for Compositional Vector Quantization
* Looking Broader for Knowledge Distillation via Receptive-Field Alignment
* LoRASculpt: Harmonious Low-Rank Adaptation for Multimodal Large Language Models
* Lorentz Entailment Cone for Semantic Segmentation
* Lose Your Self (LoYS): an adversarial entropy-based unsupervised approach for model debiasing
* LOST-3DSG: Lightweight Open-Vocabulary 3D Scene Graphs with Semantic Tracking in Dynamic Environments
* Low-Rank Expert Merging for Multi-Source Domain Adaptation in Person Re-Identification
* LVM-Lite: Training Large Vision Models with Efficient Sequential Modeling
* M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion Models
* M-FSAD-KD: Full-Link Multi-Granularity Distillation for SAR Object Detection
* M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models
* Machine-Learning-Based Radar Quantitative Precipitation Estimation Using Intra-Hour Temporal Features from Dual-Polarization Observations
* MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
* MAFM3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI
* MageBench: Bridging Large Multimodal Models to Agents
* MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation for Effective-and-Efficient Vision-and-Language Navigation
* MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
* MagicSeg: Open-World Segmentation Pretraining via Counterfactural Diffusion-Based Auto-Generation
* MANTA: Physics-Informed Generalized Underwater Object Tracking
* MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
* Mapping Vegetation Alliances Using Deep Learning and Multi-Source Remote Sensing Data
* MapVerse: A Benchmark for Geospatial Question Answering on Diverse Real-World Maps
* MARC-Net: A Modality-Availability-Aware Robust Change Network for Missing-Optical Bi-Temporal Optical-SAR Change Detection of Reclaimed Cropland
* Marine Radar Oil Spill Detection Method Based on RBM and Improved Quantum Golden Jackal Optimization Algorithm
* MarineEval: Assessing the Marine Intelligence of Vision-Language Models
* MARS: a Multimodal Alignment and Ranking System for Few-Shot Segmentation
* Marshaled Learning: Bridging Large Neural Networks with Memory-Constrained Trusted Execution Environments in Federated Learning
* Mask-Guided Self-Supervised Video Object Segmentation
* Matching Semantically Similar Non-Identical Objects
* MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
* MBTI: Metric-Based Textual Inversion for Fine-Grained Image Generation
* MDUNet: Multimodal Decoding UNet for Passive Occluder-Aided Non-line-of-sight 3D Imaging
* Mean-Shift Distillation for Diffusion Mode Seeking
* MEDAL: multi-modal MEta-space Distillation and ALignment for Visual Compatibility Learning
* MedPEFT-CL: Dual-Phase Parameter-Efficient Continual Learning with Medical Semantic Adapter and Bidirectional Memory Consolidation
* MedProbCLIP: Probabilistic Adaptation of Vision-Language Foundation Model for Reliable Radiograph-Report Retrieval
* MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
* MEGA-PCC: A Mamba-based Efficient Approach for Joint Geometry and Attribute Point Cloud Compression
* MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering
* Mem-MLP: Real-Time 3D Human Motion Generation from Sparse Inputs
* MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction
* Memoire: Learning User Personas from Gallery Tags for Personalized Photo Curation
* Memory-Augmented Representation for Efficient Event-based Visuomotor Policy Learning with Adaptive Perception and Control
* mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval
* MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
* Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes
* Meta-YOLO: Metadata-Guided Real-Time Object Detector in Aerial Imagery
* MFRA-YOLOv11: Remote Sensing Small Object Detection Algorithm Based on Multiscale Feature Extraction and Region Awareness
* MFVLR: Multi-Domain Fine-Grained Vision-Language Reconstruction for Generalizable Diffusion Face Forgery Detection and Localization
* MiAFormer: Micro Analysis Transformer for Fine-Grained Micro-Expression Understanding via Scene Flow
* Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition
* Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
* milliMamba: Specular-Aware Human Pose Estimation via Dual mmWave Radar with Multi-Frame Mamba Fusion
* Minimal Solution to the Perspective-3-Point Problem for the Camera With Unknown Focal Length and Two Degrees-of-Freedom Rotation, A
* MIST: Multilingual Incidental Dataset for Scene Text Detection
* Mitigate Catastrophic Remembering via Continual Self-Paced Dual-Knowledge Purification for Noisy Lifelong Person Re-Identification
* Mitigating Backdoor Attacks via Trigger Reconstruction and Model Hardening
* Mitigating Longitudinal Performance Degradation in Child Face Recognition Using Synthetic Data
* Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
* Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation
* MIX-based Foreground and Background Patch Augmentation Guided by Physics and Material Properties for X-ray Detection
* Mixed Diffusion for 3D Indoor Scene Synthesis
* MixER: From Cross-Modal to Mixed-Modal Visible-Infrared Re-Identification
* Mixture of Global and Local Experts With Diffusion Transformer for Controllable Face Generation
* MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
* MMCM: Multimodality-aware Metric using Clustering-based Modes for Probabilistic Human Motion Prediction
* MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
* mmWEAVER: Environment-Specific mmWave Signal Synthesis from a Photo and Activity Description
* Mobile-Oriented Video Diffusion: Enabling Text-to-Video Generation on Mobile Devices Without Retraining, Compression, or Pruning
* Modality-decoupled Transformer with Memory Learning for Video-based Visible-infrared Person Re-identification
* Model-Based Multiframe Radiometric Spatial Reconstruction for Optical Satellite Video
* Model-free Domain Adaptation for Concealed Multimodal Large-Language Models
* Modeling and Learning Multiple Hypotheses for Monocular 3D Object Detection
* MoE-Enhanced Explainable Deep Manifold Transformation for Complex Data Embedding and Visualization
* Moiré Zero: An Efficient and High-Performance Neural Architecture for Moiré Removal
* MoME: Estimating Psychological Traits from Gait with Multi-Stage Mixture of Movement Experts
* Moment-Reenacting: Inverse Motion Degradation With Cross-Shutter Guidance
* MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
* Monitoring Coastal Geomorphic Change and Sediment Transport Using Kite Aerial Photography (KAP) and Structure-from-Motion (SfM) Photogrammetry
* MonSter++: Unified Stereo Matching, Multi-View Stereo, and Real-Time Stereo With Monodepth Priors
* Monthly Trophic Dynamics of Lakes and Reservoirs in Eastern China Based on Harmonized Landsat-Sentinel Observations
* MooTrack360: A Novel Fisheye Camera Dataset for Robust Multi Dairy Cow Detection and Tracking
* More Than Memory Savings: Zeroth-Order Optimization Mitigates Forgetting in Continual Learning
* MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View Consistency
* Morphing Through Time: Diffusion-Based Bridging of Temporal Gaps for Robust Alignment in Change Detection
* MorphXAI: An Explainable Framework for Morphological Analysis of Parasites in Blood Smear Images
* MoSCo: Real-time and Efficient Text-to-Motion Synthesis via Delta Training
* Motion-Aware Graph Fusion Network for 3D Human Pose Estimation
* MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models With Temporal Distillation
* MR-Pruner: Training-free Multi-resolution Visual Token Pruning for Multi-modal Large Language Models
* MSRTrack: LLM-Powered Object Tracking with Motion and Semantic Reasoning
* Multi-Agent Diffusion Approach for MRI Anomaly Segmentation via Modality-Specific LoRA Specialization, A
* Multi-Camera Self-Calibration in Sports Motion Capture: Leveraging Human and Stick Poses
* Multi-Domain Feature Integration Based Trusted Partial Multi-View Incomplete Multi-Label Learning
* Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios
* Multi-Granularity Graph Neural Network for Satellite-Assisted Marine Environmental Field Reconstruction over Sparse Observation Grids
* Multi-Modal Soccer Scene Analysis with Masked Pre-Training
* Multi-Receptive Field Ensemble with Cross-Entropy Masking for Class Imbalance in Remote Sensing Change Detection
* Multi-Source Fusion Positioning Revisited by Drawing on Human Thinking Process
* Multi-Temporal Remote Sensing Assessment and Pareto-Based Land-Use Optimization for Balancing Carbon Storage, Ecological Value, and Economic Value in Jiangsu Province, China
* Multi-View Causal Feature Selection
* Multi-View Photometric Stereo Pipeline for Specular 3D Fruit Reconstruction, A
* Multi-view stereo with multiple projectors for oneshot entire shape scan based on Neural SDF and DSSS demultiplexing
* Multiband Spectropolarimetric Signature Analysis for Material, Object, Land Cover Class, and Collection Geometry Separability
* Multidimensional Quantification of Engineering Distresses and Secondary Periglacial Hazards Along Linear Infrastructure in the Permafrost Region of Northeast China Using UAV-LiDAR and Synchronous Visible-Light Imagery
* Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
* Multimodal Graph Representation Learning over Arbitrary Sets of Modalities
* Multimodal Medical Image Binding via Shared Text Embeddings
* Multimodal Voice Activity Projection for Multi-Party Conversations
* Multiscale Decomposition Reveals the Scale-Dependent Drivers of Complex Urban Ground Deformation in Zhengzhou, China
* MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
* MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
* MuseDance: A Diffusion-based Music-Driven Image Animation System
* MVAT: Multi-View Aware Teacher for Weakly Supervised 3D Object Detection
* MVHumanNet++: A Large-Scale Dataset of Multi-View Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization
* NAPP: Noise-Adaptive Prototype Perturbation for Few-Shot Learning
* Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent Space
* NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction
* Negation-Aware Test-Time Adaptation for Vision-Language Models
* NerVast: Compression-Efficient Scaling of Implicit Neural Video Representations via Scene-based Parameter-sharing
* NERVE: Neighbourhood & Entropy-guided Random-walk for training free open-Vocabulary sEgmentation
* Network-agnostic distortion-robust projections for wide-angle image understanding
* Neural Geometry Image-Based Representations with Optimal Transport (OT)
* NEURO-GUARD: Neuro-Symbolic Generalization and Unbiased Adaptive Routing for Diagnostics - Explainable Medical AI
* NeuroBridge: Few-Shot Cross-Modal Neuron Re-identification via Dual-Channel Deep Metric Learning
* NeuroVLM: A Contrastive Vision-Language Model for Medical Reasoning in Alzheimer's Disease Diagnosis
* New Method for Extracting Short-Term Deformation Signals from InSAR Time Series and Its Application to the Haihe River 23-7 Basin-Wide Extreme Flood Event, A
* nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis
* No MoCap Needed: Post-Training Motion Diffusion Models with Reinforcement Learning using Only Textual Prompts
* No One-Size-Fits-All Neurons: Task-Based Neurons for Artificial Neural Networks
* NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining
* Non-Aligned Reference Image Quality Assessment for Novel View Synthesis
* Non-Contact Blood Pressure Estimation from Face Videos via Physiology-Aware Contrastive Learning
* Non-Convex Joint Sparse and Low-Rank Optimization for Enhanced ISAR Imaging from Incomplete Data
* Non-Invasive 3D Gait Analysis Framework for Quantifying Psychomotor Retardation in Major Depressive Disorder, A
* Non-Local Convergence Analysis of Gradient Flow for Deep Linear Networks, A
* Nordic Skiing Dataset (NSD): A Pose Estimation Dataset for Performance Feedback in Cross Country Skiing
* Not all Blends are Equal: The BlEmoRe Dataset of Blended Emotion Expressions with Relative Salience Annotations
* Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model
* Novel Calibration Method for Networked X-Band Radar Based on Opposing RHI Scans, A
* Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis, A
* NRGMark: Localized Watermarking for Energy Transparency in Images
* NullFace: Training-Free Localized Face Anonymization
* NWCSAF High Resolution Winds (NWCSAF GEO-I HRW) Stereo AMVs over the Atlantic Ocean
* ObjectCore: Efficient Few-shot Logical Anomaly Detection using Object Representations
* ObjectMeshDeform : Towards recovering precise 3D geometry of real objects via image-guided mesh deformation of 3D generative priors
* Occlusion Boundary and Depth: Mutual Enhancement via Multi-Task Learning
* ODEt(ODEl): Shortcutting the Time and the Length in Diffusion and Flow Models for Faster Sampling
* Odo: Depth-Guided Diffusion for Identity-Preserving Body Reshaping
* Odometry and Mapping for Complex Environment Perception Under Partial-View Sensing
* OMeGa: Joint Optimization of Explicit Meshes and Gaussian Splats for Robust Scene-Level Surface Reconstruction
* OMNI-Dent: Towards an Accessible and Explainable AI Framework for Automated Dental Diagnosis
* On Applicability of Synthetic Datasets for Facial Expression Recognition
* On the Evaluation of Multimodal Large Language Models for Agricultural Image Classification Across Diverse Tasks
* On the Impact of Face Segmentation-Based Background Removal on Recognition and Morphing Attack Detection
* On-The-Fly OVD Adaptation with Flame: Few-Shot Localization Via Active Marginal-Samples Exploration
* One Model, Many Behaviors: Training-Induced Effects on Out-of-Distribution Detection
* One-Cycle Structured Pruning via Stability-Driven Subnetwork Search
* One-Shot Fine-Grained Re-Identification of Paint Marked Honey Bees using Vision Foundation Models
* One-shot Portrait Stylization via Geometric Alignment
* Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
* Open Set Face Forgery Detection via Dual-Level Evidence Collection
* Open Your Eyes to See More: Dual Perspective Contrastive Learning for Skeleton-Based Action Understanding
* OpenCowID: Zero-Shot Visual Identification of Dairy Cows
* OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
* OPFormer: Object Pose Estimation leveraging foundation model with geometric encoding
* Optimal Formation Flying for Single-Pass Multi-Baseline Across-Track Synthetic Aperture Radar Interferometry
* Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods
* Optimization-Free Style Transfer for 3D Gaussian Splats
* Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
* Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
* Optimizing Trust and Safety Regions for Text-to-Image Generation in High-Dimensional Manifold Spaces
* OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting
* ORCA: Object Recognition and Comprehension for Archiving Marine Species
* Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition
* Ordinal-Aware Multimodal Engagement Recognition for Collaborative Learning
* Orthogonal Projection Steering for Bias Mitigation and ICAO Compliance in Diffusion Transformers
* OSEG: Improving Diffusion sampling through Orthogonal Smoothed Energy Guidance
* Overcoming Fine-Grained Visual Challenges in Animal Re-Identification via Semantic Feature Alignment
* Overcoming Small Data Limitations in Video-Based Infant Respiration Estimation
* OW-Rep: Open World Object Detection with Instance Representation Learning
* PADM: A Physics-aware Diffusion Model for Attenuation Correction
* PALMS+: Modular Image-Based Floor Plan Localization Leveraging Depth Foundation Model
* PaRaChute: Pathology-Radiology Cross-Modal Fusion for Missing-Modality-Robust Survival Prediction
* Parallel Diffusion Solver via Residual Dirichlet Policy Optimization
* Partial Contrastive Learning for Partially View-Aligned Multi-View Clustering
* Patch Your Matcher: Correspondence-Aware Image-to-Image Translation Unlocks Cross-Modal Matching via Single-Modality Priors
* Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
* PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
* Paving the Way for Point Cloud Video Representation Learning Using a PDE Model
* PaW-ViT: A Patch-based Warping Vision Transformer for Robust Ear Verification
* PCBClip: Vision-Language Defect Detection Model for Low-Sample Inspection Systems
* PDV: Prompt Directional Vectors for Zero-shot Composed Image Retrieval
* PEaRL: Pathway-Enhanced Representation Learning for Gene and Pathway Expression Prediction from Histology
* Perception-Inspired Color Space Design for Photo White Balance Editing
* Perceptual Observatory Characterizing Robustness and Grounding in MLLMs, The
* Perceptually Guided 3DGS Streaming and Rendering for Mixed Reality
* Performance Analysis of BDS-3 PPP-B2b During the Satellite In-Orbit Upgrade Period
* Performance of Conformal Prediction in Capturing Aleatoric Uncertainty
* Person-in-WiFi 3D: Unified Model for 3D WiFi Perception
* Personalized Image Privacy Advisors via Federated Daisy-Chaining
* PerVL-Bench: Benchmarking Multimodal Personalization for Large Vision-Language Models
* PFIRNet: UAV-to-Satellite Cross-View Self-Localization via Continuous Probability Field Inference
* Phenology-Adaptive Rubber Plantation Mapping (PARM) Framework Coupling Sentinel-1 SAR and Optimally Selected Spectral Indices Across Heterogeneous Tropical Regions, A
* Phenology-Aware Compound Heat and Drought Events and Potential Exposure for Summer Maize in the Huang-Huai-Hai Plain, China
* Photo Dating by Facial Age Aggregation
* PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
* PhysEye: A Physics-Informed Framework for Unsupervised 3D Eye Tracking with Event Cameras
* Physically Constrained Quadratic Programming for Spectral Unmixing of Nighttime Lighting: A Case Study with Data Generation Procedure
* Physics-informed Dynamic 3D Face Reconstruction from Videos
* Physics-Informed Residual Learning for Vertical Profile Reconstruction of Atmospheric Optical Turbulence from Tethered UAV Observations over the Ngari Plateau
* Physics-Informed Spatially Variant Image Restoration for Unresolved Infrared Remote Sensing Small Targets
* PHYSPLAT: A Framework for Photorealistic Hybrid Simulation of Real and Synthetic Elements using 3D Gaussian Splatting
* Pinus pinaster Seedling Detection in Coastal Dune Plantations Using a UAS Multispectral Point Cloud and Point Transformer V3
* PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
* Pixel-Level Sentinel-2 Modeling Reveals Scale-Dependent Edge Effects and Structural Heterogeneity in Semideciduous Seasonal Forest Fragments of the Brazilian Cerrado
* PKTA: Part-oriented Knowledge Transfer and Acquisition for Non-Exemplar Lifelong Person Re-Identification
* Pmaf Loss: Probabilistic Margin-Aware Focal Loss for Robust Medical Image Classification
* PMP: Plug-and-Play Physics Momentum Prior for Contact-Stable 3D Human Pose Estimation
* Point-to-Dense Supervision Framework for SAM Based Remote Sensing Segmentation with Prototype Refinement and Structure Aware, Confidence Guided Topology Learning
* Point2Pose: A Generative Framework for 3D Human Pose Estimation with Multi-View Point Cloud Dataset
* Pointmap-Conditioned Diffusion for Consistent Novel View Synthesis
* PointNet4D: A Lightweight 4D Point Cloud Video Backbone for Online and Offline Perception in Robotic Applications
* PointSt3R: Point Tracking through 3D Grounded Correspondence
* Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
* Polymorph: Energy-Efficient Multi-Label Classification for Video Streams on Embedded Devices
* PopSign v2.0: Extending an Isolated American Sign Language Dataset to Over 360,000 Examples of 562 Concepts
* Pose-Diverse Multi-View Virtual Try-on from a Single Frontal Image via Diffusion Transformer
* PoseAdapt: Sustainable Human Pose Estimation via Continual Learning Benchmarks and Toolkit
* PoseGaussian: Pose-Driven Novel View Synthesis for Robust 3D Human Reconstruction
* PosePilot-V3D: A Web-Based Framework for 3D Hierarchical Skeleton Reconstruction from Monocular Video
* PosePilot-V3D: Interactive Exploration and Analysis of 3D Human Motion from Monocular Video
* Potential Source-to-Sink Spatial Correspondence in a Martian Analog Environment of the Qaidam Basin Based on GF-5A Hyperspectral Imagery
* Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
* Practical Framework for Surface Water Extraction from GF1/GF6 Wide-Field-View Imagery, A
* Predicting Important Photons for Energy-Efficient Single-Photon Videography
* Predicting Spatial Variability of PR Protein Concentration in White Grape Juice Using UAV Multispectral Data and Machine Learning
* Predicting Task fMRI Contrasts from Resting-State fMRI Using Sparse 3D Convolutions
* Predictive Performance and Resampling-Based Prediction Uncertainty of a Stacking Ensemble for Landslide Susceptibility Assessment in Bayi District, China
* PredMapNet: Future and Historical Reasoning for Consistent Online HD Vectorized Map Construction
* Prelaunch Assessment and Correction of Polarization Effects for HIRAS-II on the Fengyun-3 Satellite
* Preprocessing Mismatch and Input Normalisation in Transferring a Multispectral Foundation Model to Marine Surface Segmentation
* Pretraining Helps When Capacity Allows: Evidence from Ultra-Small ConvNets
* PrevMatch: Revisiting and Maximizing Temporal Knowledge in Semi-Supervised Semantic Segmentation
* Prior-Guided Lightweight Dual-Task Network for Composite Active Jamming Recognition and Time-Frequency Parameter Estimation in Radar Remote Sensing
* PRISM-CAFO: Prior-conditioned Remote-sensing Infrastructure Segmentation and Mapping for CAFOs
* PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
* Privacy-Compliant Human Data Synthesis in Images for GDPR
* Privacy-Preserving Online Federated Learning for Massive Infinite Streams
* Probabilistic Scene Graph Prompting: Uncertainty-Aware Structured Reasoning in Multimodal LLMs
* Probabilistic Spatial Completion of FEMA Special Flood Hazard Area Coverage in Louisiana Using Conditional Diffusion and Distributionally Trustworthy Explanation
* Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
* Process-Informed Satellite-Ground Fusion for Coastal Compound Humid-Heat and Photochemical Oxidant Early Warning
* Progressive Curriculum Learning and Ghost-Aware Supervision for Real-World Face Detection
* Promoting Generalization for Exact Combinatorial Solvers via Adversarial Instance Augmentation
* Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation
* Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
* PromptGAR: Flexible Promptive Group Activity Recognition
* ProSkill: Segment-Level Skill Assessment in Procedural Videos
* ProtoGMVAE: A Variational Auto-Encoder with True Gaussian Mixture Prior for Prototypical-based Self-Explainability
* PS3: Part level instance segmentation in 3D
* PSA-MIL: A Probabilistic Spatial Attention-Based Multiple Instance Learning for Whole Slide Image Classification
* PSDiffusion: Harmonized Multi-Layer Image Generation via Layout and Appearance Alignment
* Pseudo Contrastive Learning for Diagram Comprehension in Multimodal Models
* PVeRA: Probabilistic Vector-Based Random Matrix Adaptation
* Pyramidal Spectrum: Frequency-based Hierarchically Vector Quantized VAE for Videos
* Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
* QAL: A Loss for Recall-Precision Balance in 3D Reconstruction
* QC-SF: Improving Computer Vision for Airborne LiDAR Point Clouds of Boreal Forests with Quebec Simulated Forest Dataset
* QCFace: Image Quality Control for boosting Face Representation & Recognition
* QuadraNet V2: Efficient and Sustainable Training of High-Order Neural Networks with Quadratic Adaptation
* Quality-driven Adaptive Morphing Attack Detection in Operational Scenarios via Online Learning
* Quality-Driven and Diversity-Aware Sample Expansion for Robust Marine Obstacle Segmentation
* Quantifying the Limits of Segmentation Foundation Models: Modeling Challenges in Segmenting Tree-Like and Low-Contrast Objects
* Quantitative Assessment of LiDAR Availability in Smoke-Filled Tunnels Using a Degradation Scoring Algorithm
* QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks
* QuEENet: Quantum-Enhanced Expressive Network for Image Classification
* QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
* R-MMA: Enhancing Vision-Language Models with Recurrent Adapters for Few-Shot and Cross-Domain Generalization
* R3: Reconstruction, Raw, and Rain: Deraining Directly in the Bayer Domain
* RAC-Net: Interpretable Medical Small Target Segmentation Network With X-Ray Radiation Attenuation Characterization
* Radargrammetric 3D Positioning of Pseudo Corner-Reflector Scatterers in KOMPSAT-5 Stacks with Per-Target Conditioning Diagnostics
* Raising the Bar in Graph OOD Generalization: Invariant Learning Beyond Explicit Environment Modeling
* RampWatch: An In-the-Wild Dataset and Text-Guided Detection Framework for Recreational Vessels
* Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery
* RapidMV: Leveraging Spatio-Angular Latent Space for Efficient and Consistent Text-to-Multi-View Synthesis
* RAT4D: Rig and Animate Objects without Surface Templates in 4D
* RAVEN: A Rapid Agentic Vision Framework for Emergency Response in Vulnerable Settlements
* RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
* RAW-Adapter: Adapting Pre-Trained Visual Model to Camera RAW Images and a Benchmark
* REACH: Hand Pose Estimation from Room Corners
* ReactionMamba: Generating Short & Long Human Reaction Sequences
* Read and Tell - Speech Dataset for Scripted and Spontaneous Scenarios
* Real-Time Tracking of Flexible Markers in Low-Contrast Fluoroscopy Using a Deep Neural Network Trained Solely on Synthetic Data
* RealDroneVision: Dataset and Architecture Advancements for Small-Object Drone Detection
* Reamil: Reasoning- and Evidence-Aware Multiple Instance Learning for Whole-Slide Histopathology
* Reason Then Ground: Multilingual Text/Logo Grounding on Movie Posters
* ReBrain: Brain MRI Reconstruction from Sparse CT Slice via Retrieval-Augmented Diffusion
* Reciprocal Teaching: Dynamic Multi-Model Teacher-Student Learning for Multiple Noisy Annotations
* Reconstructing Realistic and Relightable Eyes
* Reconstructing Satellites in 3D From Amateur Telescope Images
* Reference-based Positive Ski Coaching using Foot Pressure and MLLMs
* Referring Change Detection in Remote Sensing Imagery
* ReFineVQA: Iterative Refinement of Video Description via Feedback Generation for Video Question Answering
* Region-Aware Latent Axis Discovery for Predictive Botulinum Toxin Facial Simulation
* RegionAligner: Bridging Ego-Exo Views for Object Correspondence via Unified Text-Visual Learning
* Regionalized Uncertainty Budget for Sea-Level Trend and Acceleration Estimates in the China Seas and Their Adjacent Oceans, A
* Reinforcement Learning-based Adaptive Control of Classifier-Free Guidance and Timestep Embeddings in Diffusion Models
* Relation DETR+: Exploring Explicit Position Relation Prior for Dense Prediction
* Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
* Reliability-Aware Adaptive Band Gating with Domain Expansion for Cross-Scene Hyperspectral Band Selection
* Reliable Uncertainty Estimation via Discriminative Feature Learning for Evidential Deep Classification
* RemEdit: Efficient Diffusion Editing with Riemannian Geometry
* REMinD: Balancing Robust Concept Unlearning and Image Quality in Diffusion Models
* ReMiX-MAE: Learning Missing-Channel Cross-Modal Representations from RGB-Only Clinical Facial Videos for Sympathetic-Mediated Pain Assessment
* Remote Sensing Forestry Similarity Convolution
* Repurposing Gait Recognition Priors for Generalizable Fine-Grained Parkinson's Disease Assessment
* Restora-Flow: Mask-Guided Image Restoration with Flow Matching
* Rethinking Latent Variable in Learned Image Compression
* Rethinking Real Image Editing: Unleashing Diverse Editing Operators via Multi-Objective Optimization
* Retrieval of Warm-Season Radar Composite Reflectivity in Sichuan by Integrating FY-4A Multi-Channel Satellite Data and DEM Topographic Information
* Retrospective Forest Volume Estimation in Southern Chile Using ALOS-PALSAR for Carbon MRV Applications
* Reverse Personalization
* Revisiting an Old Perspective Projection for Monocular 3D Morphable Models Regression
* Revisiting InternVL: A Systematic Technical Framework for Building Powerful Open-Source Vision-Language Models
* Revisiting Layer Normalization for Point Cloud Test Time Adaptation
* Revisiting Retentive Networks for Fast Range-View 3D LiDAR Semantic Segmentation
* Revisiting Vision-Language Foundations for No-Reference Image Quality Assessment
* Reviving Unsupervised Optical Flow: Concept Reevaluation, Multi-Scale Advances and Full Open-Source Release
* RGB-TViT: Multimodal RGB-Thermal Transformers for Early Prediction of Vasovagal Reactions in Blood Donation
* RIFD-DETR: Rotation-Invariant Face Detection with DETR and Polar-Aware Landmarks
* RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
* RoadMark-AWAConv: Adaptive Weight-Anchor Convolution for Fine-Grained Semantic Segmentation of Road Marking Point Clouds
* Roadside Monocular 3D Detection Prompted by 2D Detection
* RobuMTL: Enhancing Multi-Task Learning Robustness Against Weather Conditions
* Robust 3D Semantic Occupancy Prediction With Calibration-Free Spatial Transformation
* Robust Model Fitting via Motion-Aware Pyramid Transformer-Guided Preference Filtering and Consensus Smoothing
* Robust Multimodal Emotion Recognition from Incomplete Modalities via Query-Based Unimodal and Cross-Modal Learning
* Robust Object Detection for Long-Term Thermal Surveillance via Explicit Background Modeling
* Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors
* Robust Thermal Image Object Detection Challenge: Advancing Multi-Object Detection Performance Under Long-Term Thermal Drift
* Robust Thermal Image Object Detection Via Appearance-Guided Mixture of Experts
* RobustFormer: Noise-Robust Pre-training for Images and Videos
* RobustGait: Robustness Analysis for Appearance Based Gait Recognition
* Robustness of SAM: Segment Anything Under Corruptions and Beyond
* Role of Context in Prosocial Affect Recognition, The
* Role of Language-Guidance in Knowledge Distillation for Semantic Segmentation Under Limited Field-Of-View Autonomous Driving
* RoLID-11K: A Dashcam Dataset for Small-Object Roadside Litter Detection
* Root Completion from Intraoral Scans of Tooth Crowns using Diffusion with Patch Perturbation
* RPCANet^++: Deep Interpretable Robust PCA for Sparse Object Segmentation
* RPT-SR: Regional Prior attention Transformer for infrared image Super-Resolution
* S2O: Static to Openable Enhancement for Articulated 3D Objects
* SaccadeX: Directed Acyclic Graph-based Semi-Supervised Learning of Continuous Ocular Dynamics from Sparse Neuromorphic Streams
* Safe Image Authenticity Challenge: Detecting and Localizing Partial and Fully Synthetic Manipulations, The
* Safe Vision-Language Models via Unsafe Weights Manipulation
* SafeguardGS: 3D Gaussian Primitive Pruning While Avoiding Catastrophic Scene Destruction
* SAFER-AiD: Saccade-Assisted Foveal-peripheral vision Enhanced Reconstruction for Adversarial Defense
* SAIL: Self-supervised Learning of Lighting-Invariant Representations from Real Images with Latent Diffusion
* Salience-SGG: Enhancing Unbiased Scene Graph Generation with Iterative Salience Estimation
* Saliency-Guided DETR for Moment Retrieval and Highlight Detection
* SAR-Oriented and Physics-Guided Ocean Wave Spectrum Retrieval
* SasMamba: A Lightweight Structure-Aware Stride State Space Model for 3D Human Pose Estimation
* Satellite Remote Sensing for Fishing Vessel Identification and Monitoring: A Comparative Analysis of Modalities and a Review of Datasets
* Satellite Remote Sensing of a Melting Glacier Albedo: Examples from EnMAP and an Intercomparison with Other Satellite and Ground Measurements
* Satellite-Based Detection of Looted Archaeological Sites Using Machine Learning
* Satellite-Derived Shorelines Reveal Typhoon-Driven Erosion and Monsoon-Gated Recovery on the Macrotidal Coast of Fujian, China
* Satellite-Driven Spatiotemporal Multiscale Perception Learning for Estimating Daily Arctic Sea Ice Thickness
* Satellite-UAV Collaborative Off-Road Traversability Mapping and Incremental Updating for Unmanned Ground Vehicles
* SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
* SAVeD: Learning to Denoise Low-SNR Video for Improved Downstream Performance
* Savior: Sample-Efficient Adaptation of Vision-Language Models for OCR Representation
* SCAdapter: Content-Style Disentanglement for Diffusion Style Transfer
* Scalable Video Action Anticipation with Cross Linear Attentive Memory
* Scale-Aware Gated Routing Distillation for Lightweight Remote Sensing Object Counting
* SCALEX: Scalable Concept and Latent Exploration for Diffusion Models
* Scaling up Occupancy-Centric Driving Scene Generation: Dataset and Method
* Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination
* Scanpath Prediction in Panoramic Videos via Expected Code Length Minimization
* SCATR: Mitigating New Instance Suppression in LiDAR-based Tracking-by-Attention via Second Chance Assignment and Track Query Dropout
* ScatSpotter - A Dog Poop Detection Dataset
* Scattering-Aware Latent Field Modulation for Synthetic Aperture Radar Object Detection
* SceneEdited: A City-Scale Benchmark for 3D HD Map Updating via Image-Guided Change Detection
* SceneEval: Evaluating Semantic Coherence in Text-Conditioned 3D Indoor Scene Synthesis
* SceneProp: Combining Neural Network and Markov Random Field for Scene-Graph Grounding
* SceneShine: Illumination-aware Human Scene Gaussian Re-Splatting from Mobile Device Video
* ScoliGaitX: A Deep Multi-Modal Fusion Network for Scoliosis Assessment via Gait Video Analysis
* SCORE: Soft Label Compression-Centric Dataset Condensation via Coding Rate Optimization
* ScoreNet: Netting Lightweight Quality Scores for Better Visual Assessment with Large Multi-Modality Models
* SCORP: Scene-Consistent Object Refinement via Proxy Generation and Tuning
* Screening for Relative Risk of Low Soil Fertility in Mown-Grazed Grasslands of the Qinghai-Tibet Plateau Using Multi-Year Hydrothermal Backgrounds
* SD-CSFL: A Synthetic Data-Driven Conformity Scoring Framework for Robust Federated Learning
* SDPT: Synchronous Dual Prompt Tuning for Visual-Language Pre-Trained Models
* SDT-6D: Fully Sparse Depth-Transformer for Staged End-to-End 6D Pose Estimation in Industrial Multi-View Bin Picking
* Sea-CLIP: Mining Semantic-Aware Representations for Few-Shot Anomaly Detection with CLIP
* SeaClips: A Video Dataset for Maritime Object Detection
* Seafloor Morphology and Inner Shelf Benthic Habitats of the Sinuessa Shallow Coralligenous Bank, Eastern Tyrrhenian Margin
* SeaScope: A Transparent and Reproducible LLM-Assisted Framework for Maritime Earth Observation Analysis
* Seasonal Consistency Between Solar-Induced Chlorophyll Fluorescence and Vegetation Indices Across Global Urban Ecosystems
* See, Record, Do: Automated Generation of UI Workflows from Tutorial Videos
* See, Think, Learn: A Self-Taught Multimodal Reasoner
* Seeing in the Dark: Synthesizing Underexposure for More Robust Underwater Image Augmentation
* Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
* Seeing Isn't Believing: Context-Aware Adversarial Patch Synthesis via Conditional GAN
* SEEKr: Efficient Knowledge Distillation for Face Recognition
* SegMango: Early Deep Mango Yield Prediction based on Flower Segmentation and Weather Data
* Segment Anything but Farms: Comparing Segmentation Paradigms for Rural UAV Captured Ultra-High-Resolution Imagery
* Segmentation-Aware Latent Diffusion for Satellite Image Super-Resolution: Enabling Smallholder Farm Boundary Delineation
* SegMo: Segment-aligned Text to 3D Human Motion Generation
* Selecting and Distilling Cross-Label Models
* Self-Supervised Compression and Artifact Correction for Streaming Underwater Imaging Sonar
* Self-Supervised Visual Prompting for Cross-Domain Road Damage Detection
* Semantic Map Guided Bird's-Eye View Learning for Online HD Map Construction
* Semantic Segmentation for 3D Point Clouds with Curvature-Aware Sampling and Inverse-Density Weighting
* Semantic-Texture Complementation and Prediction-Guided SAM Fusion for Plastic Mulch Segmentation in GF-7 Imagery
* Semi Synthetic Iris Image Generation with Identity Preservation Using Latent Diffusion Models
* Semi-supervised Domain Adaptation via Mutual Alignment through Joint Error
* Semi-Supervised Hierarchical Open-Set Classification
* Semi-supervised Key-Point Estimation for Echocardiography Video
* SENCA-st: Integrating Spatial Transcriptomics and Histopathology with Cross Attention Shared Encoder for Region Identification in Cancer Pathology
* SeqFeedNet: Sequential Feature Feedback Network for Background Subtraction
* SeqPE: Transformer With Sequential Position Encoding
* Severe Positive Ionospheric Storm at American Low Latitudes During an Intense Long-Lasting CEJ Period of the August 2018 Geomagnetic Storm
* SFMNet: Sparse Focal Modulation for 3D Object Detection
* SFPRNet: A Spatio-Frequency Synergistic Progressive Restoration Network for Infrared Image Destriping
* SGD-Mix: Enhancing Domain-Specific Image Classification with Label-Preserving Data Augmentation
* SGFormer: Simplifying and Scaling Graph Transformers With Single-Layer Attention and Approximation-Free Linear Complexity
* SGPMIL: Sparse Gaussian Process Multiple Instance Learning
* SGW-GAN: Sliced Gromov-Wasserstein Guided GANs for Retinal Fundus Image Enhancement
* Shades of Generalization: Diversity Aware Open Set Deepfake Detection
* Shape vs. Texture: Influence of Geometric Structure and Surface Patterns on Face Recognition
* Shape-Based Object Detection via Gesture Prompts: Leveraging Pre-trained Open-Vocabulary Models
* SHaSaM: Submodular Hard Sample Mining for Fair Facial Attribute Recognition
* Shift-Equivariant Complex-Valued Convolutional Neural Networks
* Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
* SIAM: Synchronous Interaction Attention for Human Mesh Recovery
* SiamNet: A Double-Temporal SAR Avalanche Detection Method Integrating Multiscale Features and an Attention Mechanism
* SiamTrackCaps: A Siamese-Based Single Object Tracker Using Capsule Networks
* SignMoD: Sign Language Video Generation via Mixture of Diffusion
* SilverLining: Data-First Mitigation of Spatial and Spectral Shortcuts Without Introducing New Confounders
* SimForce: Force and Surface Electromyography from Full Body Video with Graph Neural Nets
* Similarity-aware Probabilistic Embeddings Modeling for Video-Text Retrieval
* Simplifying Spatio-Temporal Graphs in Skeleton-Based Gait Analysis
* Single-Stage Instance Segmentation Survey: A 1+2+3 Technical Framework
* Single-step Diffusion for Image Compression at Ultra-Low Bitrates
* Single-Step Latent Diffusion for Underwater Image Restoration
* SJD++: Improved Speculative Jacobi Decoding for Training-Free Acceleration of Discrete Auto-Regressive Text-to-Image Generation
* Skarimva: Skeleton-based Action Recognition is a Multi-view Application
* SkelSplat: Robust Multi-view 3D Human Pose Estimation with Differentiable Gaussian Rendering
* Sketch-guided Cage-based 3D Gaussian Splatting Deformation
* Sketch2Stitch: GANs for Abstract Sketch-Based Dress Synthesis
* Sketch3R: Rapid and Realistic 3D VR Sketch Creation to Shape Retrieval
* SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
* SmoothDiffusion-VE: Real-time Generative Video Editing Using Adaptive Feature Cache
* Snapmoji: Instant Generation of Animatable Dual-Stylized Avatars
* SOAF: Scene Occlusion-aware Neural Acoustic Field
* SODA: Semantic-Cohesive Knowledge Distillation for Deep Cross-modal Hashing
* Soil Moisture Retrieval Based on Multi-Temporal Dual-Polarization Brightness Temperature Parameterization
* SOLAR: Switchable Output Layer for Accuracy and Robustness in Once-for-All Training
* SOPHY: Generating Simulation-Ready Objects with PHYsical Materials
* SortWaste: A Densely Annotated Dataset for Object Detection in Industrial Waste Sorting
* SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models
* SpaceEra++: A Unified Framework Toward 3D Spatial Reasoning in Video
* Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains
* SPAR-Det: Segmentation-guided and Prior-Aided Routing for Small Object Detection
* Spatial Domain Dependence Evolution of Input Parameter Importance in Soil Moisture Retrieval Under the XGBoost and SHAP Framework
* Spatial Prompt and Wavelet Mamba-Based Multi-Scale Cross-Domain Feature Fusion Network for Segmentation of Mining-Disturbed Land
* Spatially Complete Monthly XCH4 Mapping over China from 2019 to 2024: A Multi-Model Analysis Using Machine Learning
* Spatio-Temporal Dynamics of Mining-Induced Surface Disturbance and Backfilling in Open-Pit Coal Mines Across China's Arid and Desert Regions (1990-2023)
* Spatiotemporal Deep Learning for Continuous Illumination Mapping and Sun-Synchronous Path Planning in Lunar Polar Exploration Under Chang'E-7 Mission Constraints
* Spatiotemporal Dynamics and Climatic Responses of Rubber Plantations' Aboveground Biomass in Western Hainan Island Based on Multi-Source Remote Sensing and Explainable Machine Learning
* Spatiotemporal Evolution and Multi-Factor Driving Mechanism of Land Subsidence in Shanghai Hongqiao Transport Hub Core Area Based on SBAS-InSAR (2015-2024)
* Spatiotemporal Variations and Associated Environmental Factors of Coastal Polynyas in the Kara-Laptev Seas from 2003 to 2025
* Spatiotemporal Variations in Aerosol Optical Depth and Their Relationships with Cloud Properties and Precipitation over Sudan: Insights from Satellite Observations and CMIP6 Model Projections
* Spec-Gloss Surfels and Normal-Diffuse Priors for Relightable Glossy Objects
* SpecGen: Neural Spectral BRDF Generation via Spectral-Spatial Tri-plane Aggregation
* Spectral-Adaptive Modulation Networks for Visual Perception
* Spectral-Spatial Decoupling and Fusion Network with Multi-Scale Perception for Multispectral Image Compression, A
* SphereEdit: Spherical Semantic Editing in Diffusion Models
* SpikeRain: Towards Energy-Efficient Single Image Deraining with Spiking Neural Networks
* Splannequin: Freezing Monocular Mannequin-Challenge Footage with Dual-Detection Splatting
* Splatter Layout: Geometry-embedded 3D Reconstruction via Surface Unfolding
* SPOC: Spatially-Progressing Object State Change Segmentation in Video
* SPoRC-VIST: A Benchmark for Evaluating Generative Natural Narrative in Vision-Language Models
* SSMRadNet: A Sample-wise State-Space Framework for Efficient and Ultra-Light Radar Segmentation and Object Detection
* SSMT-Net: A Semi-Supervised Multitask Transformer-Based Network for Thyroid Nodule Segmentation in Ultrasound Images
* SSplain: Sparse and Smooth Explainer for Retinopathy of Prematurity Classification
* ST-PaveCLIP: A Spatio-Temporal Vision-Language Framework for Road Anomaly Segmentation in Images and Videos
* ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
* Stabilizing Direct Training of Spiking Neural Networks: Membrane Potential Initialization and Threshold-robust Surrogate Gradient
* Stabilizing Intrinsic Explanations: A Geometric Perspective on Concept Leakage
* STAMP-GAN: A Spatiotemporal Attention-Modulated Generative Adversarial Network for Precipitation Nowcasting
* Standardized Benchmark for Skeleton-Based Rehabilitation Assessment Using Deep Learning, A
* STARS: Self-supervised Tuning for 3D Action Recognition in Skeleton Sequences
* START: Spatial and Textual Learning for Chart Understanding
* Stationary (and Therefore Compatible) Representation is All You Need, A
* STEC: A Spatio-Temporal Entropy Coverage Metric for Evaluating Sampled Video Frames
* STEG-AIW: Spatio-Temporal Gating and Adaptive-Timestep Inference for Efficient Spiking Neural Networks
* STEN-FAWA: Spatial-Temporal Expert Network with Forgery-Aware Weighted-Adaptive Aggregator for Deepfake Detection
* Stochastic Approximation Approaches to Group Distributionally Robust Optimization and Beyond
* Streaming Real-Time Trajectory Prediction Using Endpoint-Aware Modeling
* StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
* STRinGS: Selective Text Refinement in Gaussian Splatting
* Stroke Modeling Enables Vectorized Character Generation with Large Vectorized Glyph Model
* Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
* Structured Analysis and Taxonomy of Scene Graph Representations for Group Activity Understanding, A
* Structured Context Learning for Generic Event Boundary Detection
* Structured Light With a Million Light Planes Per Second
* Style-Friendly SNR Sampler for Style-Driven Generation
* StyleDiT: A Unified Framework for Diverse Child and Partner Faces Synthesis with Style Latent Diffusion Transformer
* Submerged Hazard Identification and Processing Using Augmented Image-Based Detection (SHIP-AID)
* Subspace-Guided Knowledge Distillation for Efficient Model Transfer
* Subtle Motion Blur Detection and Segmentation from Static Image Artworks
* SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities
* Sun-E: Dataset and Benchmark for Event-Based Sun Sensing
* SuperRivolution: Fine-Scale Rivers from Coarse Temporal Satellite Imagery
* Supporting Ultra-High-Resolution Digital Agriculture Tasks with Fully Synthetic Curriculum Learning
* Surface Subsidence Analysis and Prediction in an Open-Pit Mine Using Time-Series InSAR and a CL-TSF Hybrid Model
* Surface Thermal State, Antecedent Hydroclimate, and Post-Fire Vegetation-Water Response in the Zambezi River Basin: A Multi-Source Environmental Time-Series Analysis
* SurfDist: Interpretable Three-Dimensional Instance Segmentation Using Curved Surface Patches
* Surgical Gaussian Surfels: Highly Accurate Real-time Surgical Scene Rendering using Gaussian Surfels
* SurgXBench: Explainable Vision-Language Model Benchmark for Surgery
* Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications, A
* Survey on Autonomy-Induced Security Risks in Large Model-Based Agents, A
* SVD-Det: A Lightweight Framework for Video Forgery Detection Using Semantic and Visual Defect Cues
* SVS-GAN for Semantic Synthesis of Traffic Videos for Autonomous Driving
* SWH Retrieval from SWOT KaRIn Data by Combining Backscattering and Interference Characteristics
* SWIFT: A Small-World Interaction Framework for Flow-Aware Trajectory Prediction in Autonomous Driving
* SymNet: A Multi-Task Network for Joint Radio Map Reconstruction and Transmitter Localization
* SynCAF: Synthetic Counterfactual Activations for Fairness in Facial Cognition Signals
* SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
* SynPlay: Large-Scale Synthetic Human Data with Real-World Diversity for Aerial-View Perception
* SynSacc: A Blender-to-V2E Pipeline for Synthetic Neuromorphic Eye-Movement Data and Sim-to-Real Spiking Model Training
* Synthesizing Compositional Videos from Text Description
* Synthetic Priors for Real-World Detection: A Label-Free Framework for Identifying Ultra-Rare Objects
* SynthForm: Towards a DLA-Free E2E Form Understanding Model
* Systematic Analysis of the Unintentional CSAM-Generation-Potential of Text-to-Image Models
* S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix
* T2LF: LLM-Guided Multimodal Diffusion for Text-to-Light Field Synthesis
* T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
* TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
* Tables Decoded: DELTA for Structure, TarQA for Understanding
* Tables Guide Vision: Learning to See the Heart through Tabular Data
* TacticalCalib: End-to-End 6-DoF Camera Pose Regression for Tactical Camera Calibration
* TalkingHeadBench: A Multi-Modal Benchmark & Analysis of Talking-Head DeepFake Detection
* TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
* TandemNet: A Multi-Scale Multiple-Instance Learning Framework for Early-Season Rice Yield Prediction
* TAPO: Task-Decoupled Alignment and Pseudo-Label Optimization for Cross-Scene Inshore SAR Ship Detection
* Target Superresolution Reconstruction Approach for Bistatic Airborne Radar Based on Joint Convolution Echo Model
* Task-KV: Task-Aware KV Cache Optimization via Semantic Differentiation of Attention Heads
* Task-Tailored Pre-Processing: Fair Downstream Supervised Learning
* TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning
* TED-4DGS: Temporally Activated and Embedding-based Deformation for 4DGS Compression
* Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
* Terrain Effects on Time-Frequency Characteristics of Negative Return Stroke Electric Fields: A Path-Incremental Method
* Terrain-Corrected Vegetation Index Strategy for Improving Leaf Area Index Estimation in Mountainous Areas, A
* Test Time Adaptation Using Adaptive Quantile Recalibration
* Test-time Adaptation for 3D Human Pose and Shape Estimation
* Test-Time Adaptation for Video Highlight Detection Using Meta-Auxiliary Learning and Cross-Modality Hallucinations
* Test-Time Adaptation through Semantically-guided Feature Decomposition for Few-shot Chest X-ray Diagnosis
* Test-Time Consistency in Vision Language Models
* Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
* TFCRNet: Dual-Discriminator SAR-to-Optical Translation and Region-Gated Cross-Attention Fusion for Thick-Cloud Removal
* TFFM: Topology-Aware Feature Fusion Module Via Latent Graph Reasoning for Retinal Vessel Segmentation
* Theoretical Characterization of the Good Properties of Extremely Randomized Trees for Random Forest-Distance Computation, A
* Three-Dimensional Displacement Analysis and Statistical Modeling of the Pubugou Rockfill Dam Using Multi-Track InSAR
* TICLS: Tightly Coupled Language Text Spotter
* TimeRefine: Temporal Grounding with Time Refining Video LLM
* Timestamp Query Transformer for Temporal Action Segmentation
* TM-Adapter: Temporal Merge Adapter for Efficient Global Temporal Modeling
* Tomographic Sparse View Selection Using the View Covariance Loss
* TopoRec: Point Cloud Recognition Using Topological Data Analysis
* Toward a Cross-Domain Taxonomy of Motion Quality Metrics
* Towards Consistent and Efficient Decision-Based Attacks
* Towards Egocentric 3D Hand Pose Estimation in Unseen Domains
* Towards Facilitated Fairness Assessment of AI-Based Skin Lesion Classifiers Through GenAI-Based Image Synthesis
* Towards Fast and Scalable Normal Integration using Continuous Components
* Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score
* Towards High-Fidelity, Identity-Preserving Real-Time Makeup Transfer: Decoupling Style Generation
* Towards Inclusive Biometrics: Synthetic Generation of Vitiligo Faces and Their Impact on Face Image Quality
* Towards Lightweight and Accurate Remote-Sensing Image Super-Resolution via Reparameterized Feature Enhancement Network
* Towards Pareto Efficiency in Fair Facial Expression and Action Unit Recognition
* Towards Photorealistic Style Transfer with Multimodal Guidance and Robustness to Content Images in Arbitrary Styles
* Towards Reliable Test-Time Adaptation: Style Invariance as a Correctness Likelihood
* Towards Streaming LiDAR Object Detection with Point Clouds as Egocentric Sequences
* Towards Trustworthy Face Avatars: Dynamic 3D Reconstruction with Faithful Eye Motion Modeling
* Towards Unconstrained Cross-View Pose Estimation
* Towards Unconstrained Human-Object Interaction
* TRACE: Confounder-free Adversarial Fine-tuning for Robust Object Detection
* Tracing 80 Years of Forest Succession and Species Trajectories Using Historical Aerial Photography and Sentinel-2 Imagery
* TrafficRAG: Temporal Grounding for Traffic Violations via Retrieval-Augmented Generation
* TraGraph-GS: Trajectory Graph-based Gaussian Splatting for Arbitrary Large-Scale Scene Rendering
* Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
* Training-free Detection of Text-to-video Generations via Over-coherence
* Training-Free Few-Shot Segmentation via Vision-Language Guided Prompting
* Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting
* Training-free Multi-view 4D Human Motion Reconstruction Virtual Reality System
* Training-Free Semantic Multi-Object Tracking with Vision-Language Models
* Training-Free Target Emphasis With Sam2 Pseudo-Masks for Robust Single Object Tracking
* Trajectory Tactics: When Transformers Learn Exploration to Generate Online Signature
* Trajectory-Guided Photon Accumulation for Photon-Efficient LiDAR Remote Sensing in Low-SBR Dynamic Scenes
* TranSegNet: A Transformer Guided Framework for Enhanced Segmentation of Infarct Core and Affected ASPECTS Regions for Ischemic Stroke Assessment
* TransFIRA: Transfer Learning for Face Image Recognizability Assessment
* Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups
* Transformers With Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
* Transforming Video Subjective Testing with Training, Engagement, and Real-Time Feedback
* TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
* TriaGS: Differentiable Triangulation-Guided Geometric Consistency for 3D Gaussian Splatting
* Trust-Guided Multimodal LLM Integration with Reinforcement Learning for Autonomous Driving
* TS-PCI: Point Cloud Frame Interpolation with Time-Aware Point Cloud Sampling and Self-Supervised Learning Strategy
* Two-Stage MAE with Dual-Asymmetry Learning for 3D Facial Paralysis Grading
* Type-Aware Ranking of Urban Similarity From Aerial Imagery
* UAV Applications in Forest Regeneration Survey: A Review and Case Study
* UAV Visual Localization Method Based on Token-Level Local Matching Reranking and Neighborhood-Consistent Position Fusion
* UCDSC: Open Set UnCertainty aware Deep Simplex Classifier for Medical Image Datasets
* UEOF: A Benchmark Dataset for Underwater Event-Based Optical Flow
* UI-Styler: Ultrasound Image Style Transfer with Class-Aware Prompts for Cross-Device Diagnosis Using a Frozen Black-Box Inference Network
* UltraClean: A Simple Framework to Train Robust Neural Networks against Backdoor Attacks
* Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
* Uncertainty-Aware Vision-Language Segmentation for Medical Imaging
* Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models
* Understanding Generative AI Capabilities in Everyday Image Editing Tasks
* Understanding Human-Like Biases in VLMs via Subjective Face Analytics
* Understanding the Visual Projection Space of Multimodal LLMs
* UnderWater SLAM with Laser-light sectioning method using ST-GAT
* Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
* UniCalib: Targetless LiDAR-camera Calibration via Probabilistic Flow on Unified Depth Representations
* UniCoRN: Latent Diffusion-based Unified Controllable Image Restoration Network across Multiple Degradations
* UniDiff: Parameter-Efficient Adaptation of Diffusion Models for Land Cover Classification with Multi-Modal Remotely Sensed Imagery and Sparse Annotations
* Unified Alignment Protocol: Making Sense of the Unlabeled Data in New Domains
* Unified Control for Inference-Time Guidance of Denoising Diffusion Models
* Unified Diffusion-Based Framework for Multi-Agent Trajectory Prediction Integrating Structured Multi-Modal Representations, A
* Unified Framework for Individual Tree Segmentation and Forest Biometrics Derivation from LiDAR Point Clouds Captured by Different Platforms in Diverse Forest Environments, A
* Unified Framework for Pseudo-Supervised Clustering via Weighted Sample Aggregation, A
* Unified Video Anomaly Detection Model for Detecting Different Anomaly Types
* UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training
* UniMM: A Unified Mixture Model Framework for Multi-Agent Simulation
* UniTabBank: A Large Scale Multi-Lingual, Multi-Layout, Multi-Type, Multi-Format Dataset for Table Detection
* Universal Neural Architecture Space: Covering ConvNets, Transformers and Everything in Between
* Universal Self-Attention Enhancement for Bridging Low-bit Quantization and Vision Transformers, A
* UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
* Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
* UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
* Unsupervised 3D Human Pose Estimation via Conditional Multi-view Ancestral Sampling
* Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
* Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
* Unsupervised Modular Adaptive Region Growing and RegionMix Classification for Wind Turbine Segmentation
* Unsupervised Segmentation by Diffusing, Walking and Cutting
* Unsupervised Spatially Aware Gaussian Mixture Model via Implicit Deep Priors
* Uplifting 2D to 3D Human Poses with Joint Rotations and Bone Constraints: A Strong Baseline for Sports and Fitness Applications
* Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation
* V2XScene: Multi-View Consistent 3D Scene Simulation for Collaborative Perception
* VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
* VAOT: Vessel-Aware Optimal Transport for Retinal Fundus Enhancement
* Variational Mean-Field Control Framework for Graph Representation Learning, A
* VAST-ReID: A Low-Light Benchmark Dataset for Person Re-Identification with Visual and Attribute-Rich Semantic Tracking
* VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics
* Vegetation Mapping Through Multiscale Remote Sensing
* VFace: A Training-Free Approach for Diffusion-Based Video Face Swapping
* VIBEFACE - Video and Image Biometric Dataset for Evaluation of Faces
* Video and Language Alignment in 2D Systems for 3D Multi-object Scenes with Multi-Information Derivative-Free Control
* VideoForge: Efficient Domain Adaptation for Video Generation Through Quality-Driven Rewards and Enhanced LoRA
* VideoSketcher: A Training-Free Approach for Coherent Video Sketch Transfer
* View-aware Cross-modal Distillation for Multi-view Action Recognition
* ViGG: Robust RGB-D Point Cloud Registration using Visual-Geometric Mutual Guidance
* Virtually Unrolling the Herculaneum Papyri by Diffeomorphic Spiral Fitting
* Visibility guided Self-Supervised Occlusion-Resilient Human Pose Estimation
* Visible Nearshore Object Detection in Overhead Surveillance Imagery: A Large-Scale Dataset and Benchmark
* Vision Language Models Learn to Assess Images with Specialists
* Vision Transformers for Face Recognition Need More Registers
* Vision-Informed Semantic Text Alignment for Open-set Recognition in Remote Sensing
* Vision-Language Temporal Analysis for Illegal Waste Dumping Detection in Surveillance Videos
* VISTA: A Vision and Intent-Aware Social Attention Framework for Multi-Agent Trajectory Prediction
* ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion Models
* Visual Detector Compression via Location-Aware Discriminant Analysis
* ViT-FREE: Efficient Face Recognition via Early Exiting and Synthetic Adaptation
* VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models
* ViTNT-FIQA: Training-Free Face Image Quality Assessment With Vision Transformers
* VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
* VIZOR: Viewpoint-Invariant Zero-Shot Scene Graph Generation for 3D Scene Reasoning
* VLA4CoDrive: Vision-Language-Action Dataset for Cooperative Autonomous Driving
* VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
* VLMs Guided Interpretable Decision Making for Autonomous Driving*
* VOCAL: Visual Odometry via ContrAstive Learning
* Volumetric Impact Characterization of the 2025 Palisades and Eaton Fires Using Aerial LiDAR
* VRAgent: Self-Refining Agent for Zero-Shot Multimodal Video Retrieval
* WALDO: Where Unseen Model-based 6D Pose Estimation Meets Occlusion
* WarpRF: Multi-View Consistency for Training-Free Uncertainty Quantification and Applications in Radiance Fields
* We Still See Broken Limbs: Towards Anatomical Realism in GenAI Via Human Preference Learning
* What Drives the Glacier Retreat, and How Do We See It? A Study of Measurement Methods and Environmental Drivers of Retreat in the Amundsenisen Glacial System, Svalbard
* What Happens When: Learning Temporal Orders of Events in Videos
* What Matters in Reinforcement Learning Based Training for Medical Vision Language Model: an Empirical Study
* When AI Watches the Dose: Can Vision-Language Models Perform TB Medication Adherence Assessments Like Clinicians?
* When Probe and Gallery Are Low Quality: Decreasing Accuracy and Increasing Demographic Disparities in 1:N Identification
* Where is the Watermark? Interpretable Watermark Detection at the Block Level
* WiSAR3D - Aerial LiDAR dataset for 3D object detection
* Wisdom is Knowledge Combined With Intellect: Knowledge-Embedded Hypergraph-of-Thought Reasoning for Visual Abductive Ratiocination
* WiSE-OD: Benchmarking Robustness in Infrared Object Detection
* Woman with a Knife or A Knife with a Woman? Measuring Directional Bias Amplification in Image Captions, A
* WorkZone3D: A Multimodal Dataset for 3D Work Zone Perception in Autonomous Driving
* WSSSP-Net: Weakly Supervised Semantic Segmentation Plugin Network for Face Anti-Spoofing
* WWE-UIE: A Wavelet & White Balance Efficient Network for Underwater Image Enhancement
* X-JEPA: A Novel Joint Learning Cross-Modal Predictive Alignment Framework for Remote Sensing Image Retrieval
* X3D-based Illegal Waste Dumping Detection with Temporal Localization
* XOV-Action: Toward Generalizable Open-Vocabulary Action Recognition
* YOLO-OSA: A ShuffleAttention-Enhanced YOLO Model for FOD Detection with Comprehensive Benchmarking on MS COCO
* YOLO-ROSS: A Robust Small Object Detection Model for UAV Aerial Imagery in Complex Interference Environments
* You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
* ZebraPose: Zebra Detection and Pose Estimation using only Synthetic Data
* Zero-LEAD: Source-Free Universal Domain Adaptation for Abdominal Multi-Organ Segmentation
* Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
* Zero-Shot Coreset Selection via Iterative Subspace Sampling
* Zero-Shot Domain Generalisation via Prompt-Driven Feature Refinement
* Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
* Zero-Shot Table Extraction in Business Documents: A Unified Benchmark with Error Taxonomy and Ecological Analysis
* Zero-Shot Temporal Action Localization Through Textual Guidance
* Zero-Shot Video Deraining with Video Diffusion Models
* ZonUI-3B: Competitive GUI Grounding with a 3B VLM Trained on a Single Consumer GPU
1526 for 2609

Index for "2"


Last update:21-Sep-26 19:26:29
Use price@usc.edu for comments.