Interactive segmentation of 3D medical images supports quantitative analysis of anatomical structures and disease while allowing users to specify and refine their targets. Despite substantial progress by nnInteractive and VISTA3D, reliable segmentation across diverse clinical targets remains challenging, particularly for complex anatomical structures and the heterogeneous, long-tailed spectrum of pathology. We present SAMI3D-DW V1 (hereafter SAMI3D-DW), an interactive 3D segmentation model trained on Deepwise's large-scale proprietary medical image datasets. We evaluate the model under simulated user interactions on a CT/MR benchmark comprising 4,326 cases from 219 source datasets, spanning 107 anatomical and pathological categories, organized by a medical taxonomy and evaluated with a category-balanced DSC score. SAMI3D-DW achieves the highest category-macro Dice among evaluated methods in both interaction modes. With one point, it scores 0.5764 versus 0.5315 for nnInteractive, the strongest baseline, rising to 0.7771 versus 0.7494 with five points. With bounding-box initialization, the scores are 0.7130 versus 0.6530. After five corrective clicks, SAMI3D-DW reaches 0.8002 versus 0.7868, making it the only evaluated box-compatible model to exceed 0.80. For radiologists and clinicians, SAMI3D-DW enables segmentation of complex anatomical structures, including intracranial vessel trees on CT and MR angiography, with a few clicks. In a preliminary in-house comparison involving neurofibromatosis type 1 (NF1), SAMI3D-DW-assisted tumor annotation took minutes per case and approximately one-fifteenth of the time required for manual annotation, highlighting its potential to support volumetric treatment-response assessment.
Quantum machine learning is a promising paradigm for learning from limited data, a central bottleneck in domains such as medical imaging, clinical trials, and rare diseases. Quantum convolutional neural networks (QCNNs) are particularly attractive in this setting, combining a hierarchical architecture with strong inductive bias and a parameter count that grows only logarithmically with system size. Their appeal rests on the generalization bounds of Caro et al. (2022), which show that the generalization error of a quantum model scales with the number of trainable parameters rather than with the Hilbert-space dimension, placing QCNNs in a potentially sample-efficient regime. We develop a hardware-compatible QCNN with mid-circuit measurement and classical feed-forward, and show on a binary handwritten-digit task that strong test performance is achievable from as few as 10 training samples, with the generalization error decreasing as the training set grows. At a matched 45-parameter budget the QCNN learns where an equally small classical convolutional network stays at chance, although an unconstrained classical baseline with roughly 25,000 parameters remains strongest when data are plentiful. Transpiling amplitude and angle encoded circuits across image resolutions from 2x2 to 512x512 pixels then exposes the dominant scaling bottleneck: amplitude encoding stays qubit-efficient but grows extremely deep, whereas angle encoding stays shallow but becomes qubit-prohibitive. On the medically motivated BreastMNIST benchmark the QCNN does not surpass the unconstrained classical network, yet it learns consistently above chance using orders of magnitude fewer parameters. Our results indicate that for QCNNs, learning from few samples is attainable in practice, whereas scaling to realistic image data is constrained less by optimization than by data encoding and hardware execution.
Automatic segmentation of medical images could accelerate clinical and research workflows, yet existing tools are rarely compared independently on shared benchmarks, making it difficult to assess genuine progress. This is particularly consequential for limb muscle segmentation in magnetic resonance imaging (MRI), which underpins volumetric analysis, radiomics, and fat fraction quantification used as biomarkers in neuromuscular disease -- applications where segmentation errors can directly affect interpretation. Despite the availability of multiple open-source segmentation tools spanning diverse architectures, no independent head-to-head comparison has appeared in the literature. We address this gap by evaluating eight tools, six of which run independently, and the other two of which are meant to be auxiliary algorithms once muscles have been somehow initially labeled. We evaluated these tools on three-dimensional MRI volumes across three base datasets and then later against derived datasets, assessing both quantitative metrics and qualitative usability. Our results reveal substantial performance variation across methods and datasets, with many methods showing reduced accuracy on pathological cases. Domain-specific models trained on large datasets consistently outperformed foundation models and newer general-purpose architectures. Our results reveal practical trade-offs between accuracy, generalizability, and ease of use that are critical for clinical adoption, while raising concerns for deployment on underrepresented populations and pathologies.
Neural Cellular Automata (NCAs) are a new type of neural network architecture which enable accurate and robust inference at extremely small model sizes. Recently, NCAs have advanced to become interesting low-resource alternatives to convolution- and attention-based architectures for various tasks such as image analysis, synthetic image generation, and simulation. The rapid development and increased research interest necessitate a comprehensive review of the emerging technology. This review provides an overview of the fundamentals of NCAs, applications to medical imaging, as well as insights into the state of the art. We analyze recent modifications to the originally proposed NCA architecture with respect to their efficiency and accuracy. Furthermore, we review practical applications in real-world scenarios with a focus on medical image analysis, segmentation, classification, registration, depth estimation, and image synthesis. Finally, we identify several advantages of NCAs, research gaps, and conclude with an analysis of future opportunities for NCAs in medical applications in confined settings or areas that have particular demands for robustness or efficient data processing.
This study presents a physiologically detailed biomechanical model of the mouse distal forelimb that incorporates intrinsic musculature, tendon routing, and digit-level skeletal anatomy, features simplified or omitted in existing musculoskeletal models. Using high-resolution anatomical reconstruction and computational modeling, we created a physiological representation of the wrist and digits capable of simulating complex forelimb movements. The model enables simulation of coordinated distal forelimb movement and digit-level muscle behavior during grasping-related tasks. Simulations were performed for multiple tasks, including grasping, grasping with supination, wrist flexion, and digit I flexion, with analysis focused on the grasping task due to its integration of both intrinsic and extrinsic musculature. Model performance was evaluated through comparisons of marker trajectories between torque-driven reference motion and muscle-driven simulations, temporal shuffle control, and comparisons between experimentally recorded electromyography (EMG) activity and model-predicted muscle excitation profiles. The model successfully reproduced coordinated distal forelimb kinematics, demonstrated strong agreement between torque-driven and muscle-driven simulation approaches, and generated physiologically plausible muscle excitation patterns consistent with experimentally observed EMG activity during grasping-related movement. These findings establish the model as a framework for studying fine motor control, neuromuscular coordination, and movement-related impairments in mice while providing a foundation for future investigation of neurological disorders and their underlying biomechanical mechanisms.
Objectives: Quantitative assessment of infarct hypodensity on non-contrast computed tomography (NCCT), including net water uptake (NWU), requires manual or semi-manual lesion delineation, often guided by CT perfusion or diffusion-weighted MRI, limiting clinical applicability. Automated segmentation on NCCT could enable efficient biomarker extraction such as NWU but remains challenging across heterogeneous multicenter data. This study aimed to develop and externally test a domain-aware deep learning framework for ischemic stroke segmentation on NCCT and assess its suitability for NWU quantification. Materials & Methods: In this retrospective multicenter study of 801 patients from four datasets, an nnU-Net-based model was trained on NCCT scans from the University Medical Center Hamburg-Eppendorf and the Acute Ischemic Stroke Dataset. To adapt to new domains, the model was fine-tuned on target-domain subsets from Boston (n=11) and ISLES (n=75), with evaluation on held-out cases not used for fine-tuning. Automated segmentations and NWU values were compared with expert references. Results: For lesions $\geq$ 30 mL, median Dice was 0.68 (Boston) and 0.56 (ISLES). Including smaller lesions, which predominated in ISLES, median Dice was 0.54 (interquartile range [IQR] 0.30-0.70) for acute lesion segmentation (Boston dataset) and 0.20 (IQR 0.03-0.41) for NCCT lesion segmentations when compared to post-treatment infarct (primary target of the ISLES challenge). Automated NWU mean absolute error was 1.37 percentage points (SD 1.61, Boston). Conclusion: Target-domain adaptation supported NCCT-only infarct segmentation across heterogeneous external cohorts, although performance varied across domains. The approach enabled low-error NWU quantification from baseline NCCT without advanced imaging, supporting further prospective clinical evaluation.
Lacuna, an open-source Python tool for discovering cryptic binding pockets: sites that are absent or too small to detect in a protein's unbound structure and open only during conformational fluctuation. Most binding-site predictors score a single static structure, which is precisely the structure in which a cryptic site is invisible. Lacuna instead generates a conformational ensemble from any input structure, detects pockets independently in every conformer, clusters the detections into persistent sites across the ensemble, and ranks those sites with a model fitted on within-structure pairs. Ensemble generation is pluggable: normal mode analysis by default, with implicit-solvent molecular dynamics, Boltz-2 diffusion sampling, or a user-supplied ensemble as alternatives. On the designated test fold of CryptoBench, Lacuna recovers 55.6% of cryptic sites in its top five predictions with the zero-dependency default and 66.1% with an optional PLM-assisted ranker; pooling the geometric detector with an optional learned surface detector recovers 73.9% while raising the fraction of sites found from 68.5% to 86.4%, measured on the held-out fold at five conformers. It recovers 73%, 45% and 87% on the PocketMiner set, a curated set of literature apo/holo pairs, and COACH420 respectively. The default backend completes in a median of 2.6 seconds per chain on one CPU core, so ensemble-based pocket finding does not require a simulation budget. Every site carries a continuous crypticity score, and outputs are emitted as docking-ready Boltz YAML constraints, AutoDock Vina boxes, pseudoatom PDB files, and the generated conformational ensemble as a multi-model PDB. Lacuna is MIT licensed and available at https://github.com/mooreneural/lacuna and on PyPI as lacuna-pockets.
Medical image segmentation supports quantitative clinical analysis and computer-aided diagnosis. Recent methods for medical image segmentation have improved both local feature representation and volumetric context modeling. However, existing methods still strug- gle to efficiently model cross-slice relations in anisotropic volumet- ric images, limiting segmentation consistency and accuracy. This pa- per proposes G^2RA-Net, a medical image segmentation framework that combines graph-based cross-slice relation modeling with atten- tion gating. Graph-Based Slice Relationship Modeling (GSRM) cap- tures anatomical dependencies across consecutive slices by repre- senting each slice as a graph node and propagating semantic con- text through graph message passing. The Cross-Slice Attention Gate (CSAG) then selects relevant neighboring context and emphasizes target anatomical regions through attention-guided feature modula- tion. Experiments on brain MRI and abdominal CT datasets demon- strate that G^2RA-Net outperforms representative methods in seg- mentation accuracy and boundary quality. Ablation studies further validate the proposed design.
Published medical case reports serve as a crucial medical information carrier, documenting discoveries in rare diseases, diagnostic methods, and innovative treatments. Despite the wealth of clinical knowledge in millions of case reports in the public medicine literature database (PubMed), accessing relevant information efficiently is hindered by the limitations of traditional keyword-based retrieval tools on unstructured and diverse case reports. To address the above issues, we introduce a comprehensive multimodal information system for case reports integrating structured clinical summaries of patients including medical images and biomedical named entities from 52949 open-access case reports published from 2000 to 2021. The multimodal essential information is organized in a well-structured medical ontology. Also, a powerful interface for searching and browsing case reports is designed to assist junior clinicians in retrieving cases effectively and improving the identification and diagnosis of rare diseases.
Objective: We developed a data-efficient deep learning framework for three-dimensional segmentation of pathological musculoskeletal anatomy from magnetic resonance imaging (MRI) when only limited manual annotations are available. Methods: We developed a Physics-Informed Latent-Regularized U-Net (PILR-U-Net) that combines transfer learning from healthy MRI, latent-space anatomical regularization, and elasticity-based physics-informed constraints. The physics-informed loss enforces mechanical equilibrium, near-incompressibility, and spatial smoothness of predicted deformation fields. The framework was evaluated on MRI datasets from 50 participants with cerebral palsy across 15 lower-limb musculoskeletal structures using sparse manual annotations. Performance was assessed using volumetric overlap, boundary accuracy, volume error, sensitivity, and precision. Ablation experiments evaluated the individual contributions of latent-space and physics-informed regularization. Results: Our experimental results show accurate segmentations with three-dimensional Dice coefficients ranging from 0.750 to 0.943 across evaluated structures, while most structures exhibited low surface-distance errors. Predicted deformation fields maintained positive Jacobian determinants near unity and smooth strain-energy distributions. Ablation analysis showed that both regularization components improved performance, with removal of physics-informed regularization producing the largest reductions in Dice accuracy and increases in boundary error. Conclusion: PILR-U-Net enables accurate and anatomically plausible segmentation of pathological musculoskeletal MRI under sparse supervision. Significance: Incorporating anatomical priors and biomechanical constraints into deep segmentation networks may reduce dependence on extensive pathological annotations and support patient-specific musculoskeletal modeling and clinical analysis.
Understanding how individuals organize complex natural scenes into perceptual segments is essential for explaining both typical and atypical sensory processing. In Autism Spectrum Disorder (ASD), differences in visual segmentation are widely documented. However, existing paradigms often rely on simplified stimuli and do not capture the dynamic, naturalistic processes underlying real world scene segmentation. To overcome those limitations, in this work we present a novel, multimodal experimental framework. The framework combines precisely timed, trial-based perceptual measurements with electroencephalography (EEG) to examine how neural activity aligns temporally with segmentation decisions, and with eye-tracking to examine visual exploration strategies and attention allocation. To demonstrate the framework, we report basic characterization results from a pilot study. Neurotypical (NT) participants and participants with ASD aged 16 years and above viewed natural scenes and textures and indicated whether two cued image regions belonged to the same or different perceptual segments. Each trial used a different pair of cues, and responses across trials enabled the estimation of subjectively perceived segmentation maps and the associated uncertainty. Both cohorts performed the task successfully and produced interpretable segmentation maps. Analysis of reaction time distributions, gaze profiles, and trial-aligned neural responses give us a window into cohort differences and the opportunity to study individual heterogeneity. Together, these findings establish a reproducible method for studying the temporal and spatial components of natural scene segmentation and illustrate that meaningful behavioral and neural distinctions in ASD may be studied with this approach. This framework provides a methodological bridge between controlled experimental stimuli and real-world perception in both neurotypical and neurodivergent populations.
Training deep learning-based medical image segmentation models is challenging with limited curated datasets. For AGITG TOPGEAR, a gastric cancer trial, the Clinical Target Volume (CTV) is complex and defined by multiple anatomical landmarks, making upfront training data preparation difficult for an automated contour QA segmentation model. We investigate anatomical priors, derived from surrounding organ segmentations, to provide spatial context and improve TOPGEAR CTV segmentation accuracy. We also evaluate active learning, iteratively expanding the training dataset by selecting cases expected to improve performance. One hundred TOPGEAR CT scans were retrospectively analyzed. An initial set of 10 expert-contoured cases was used to train an nnU-Net model. TotalSegmentator generated a voxel-wise anatomical prior map from surrounding structures as an additional input channel. Active learning was simulated over four iterations, selecting cases by model uncertainty and segmentation performance. All models used five-fold cross-validation for an ensemble uncertainty measure. Evaluation used a hold-out testing set of 50 cases. The anatomical prior improved CTV segmentation accuracy, increasing mean Dice Similarity Coefficient (DSC) from 0.84 to 0.86. Active learning similarly improved performance to 0.86, with greatest benefit in the final round. Combining the anatomical prior with active learning achieved the highest accuracy, with a DSC of 0.87. Model uncertainty correlated with DSC, supporting its use in identifying suboptimal predictions and guiding active learning. Anatomical priors and active learning each improved CTV segmentation accuracy and generalizability, with their combination achieving the best performance, supporting integration into segmentation model development for automated contour QA in radiotherapy clinical trials.
Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematically investigate the role of text in multimodal medical image segmentation. We first analyze several commonly used fusion strategies and find that segmentation performance is largely insensitive to the choice of fusion module. To further understand modality interactions, we propose an Evidence Decoupling Decoder (EDD) based on evidential deep learning and deep supervision. EDD serves as an internal representation analysis tool that decomposes image evidence and text-modulated evidence throughout the decoding process while maintaining competitive segmentation performance. Experimental results show that the sensitivity to text perturbation varies substantially across datasets. On BUSI and BTMRI, removing text causes catastrophic performance drops, indicating strong model reliance on textual input. On ISIC and Kvasir-SEG, text exerts relatively marginal influence. We further find that text affects predictions mainly through global semantic modulation rather than independent spatial localization, and that the specific semantic components driving text sensitivity differ across datasets. These findings provide a deeper understanding of modality interaction in multimodal medical image segmentation and offer practical insights for future model design.
主要作者:Leonhard F. Feiner、Manuel Nickel、Martin Menten
本文提出一种新的方法来量化高维输出空间中的不确定性和相关性,特别适用于医学图像分割等任务。
💡 该方法结合了随机不确定性和认知不确定性,为医学图像处理提供了更可靠的预测。
Uncertainty Quantification (UQ) plays a vital role in enhancing the reliability of deep learning model predictions, especially in scenarios with high-dimensional output spaces. This paper addresses the dual nature of uncertainty -- aleatoric and epistemic -- focusing on their joint integration in high-dimensional regression tasks. For example, in applications like medical image segmentation or restoration, aleatoric uncertainty captures inherent data noise, while epistemic uncertainty quantifies the model's confidence in unfamiliar conditions. Modeling both jointly enables more reliable predictions by reflecting both unavoidable variability and knowledge gaps, whereas modeling only one limits transparency and robustness. We propose a novel approach that approximates the resulting joint uncertainty using a low-rank plus diagonal covariance structure, capturing essential output correlations while avoiding the computational burdens of full covariance matrices. Unlike prior work, our method explicitly combines aleatoric and epistemic uncertainties into a unified second-order distribution that supports robust downstream analyses like sampling and log-likelihood evaluation. We further introduce stabilization strategies for efficient training and inference, achieving superior UQ in the tasks of image inpainting, colorization, optical flow, and depth estimation.
Accurate brain tumor segmentation from magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, surgical guidance, and disease monitoring. However, developing automated segmentation models that generalize across diverse tumor characteristics, imaging protocols, acquisition sites, and patient populations remains challenging. Variations in tumor morphology and imaging distributions can substantially degrade performance outside the training domain. Consequently, improving the robustness and generalization of deep learning-based segmentation models has become a key objective in medical image analysis. To improve segmentation robustness, we propose Multi-Stage Dynamic Prompt nnU-Net, a prompt-conditioned extension of nnU-Net. Three independent dynamic prompt modules are inserted into the deepest encoder stages. Each module contains a learnable bank of ten 256-dimensional prompt vectors and uses globally pooled encoder features to generate image-specific prompt representations. These representations are projected into feature-wise scaling $(γ)$ and shifting $(β)$ parameters that modulate encoder feature maps through Feature-wise Linear Modulation (FiLM), enabling adaptive feature conditioning at multiple semantic levels. Evaluation on the BraTS GOAT validation dataset demonstrated that the proposed Multi-Stage Dynamic Prompt nnU-Net outperformed the baseline nnU Net across the majority of evaluated metrics and tumor subregions. The proposed model achieved average lesion-wise Dice scores of 76.16% (ET), 80.04% (TC), and 86.42% (WT), compared with 74.38%, 78.14% and 84.01% for the baseline model. The results demonstrate that multi-stage dynamic prompt conditioning improves segmentation accuracy and boundary delineation for brain tumor segmentation.
Segmentation of curvilinear anatomical structures in 3D medical images remains challenging due to complex topology, severe class imbalance, weak contrast, and large variations in structure morphology. While deep learning approaches for 3D curvilinear segmentation have been proposed, they are often tailored to specific anatomies or modalities, limiting generalization across clinical settings and leaving room for improvement. Recent generative models have shown the benefits of iterative prediction for structured segmentation tasks, yet diffusion-based methods suffer from computationally expensive sampling, hindering their use on high-resolution 3D volumes. We present 3D-CurvSegFlow, a flow matching-based model for 3D curvilinear structure segmentation. The model learns a continuous transformation from a simple source distribution to the target vascular representation, enabling progressive refinement of complex curvilinear geometries with efficient inference. We evaluate our method on Three public challenging datasets covering distinct anatomies and modalities: portal vein, cerebral vessel, and coronary arteries. Using a common architecture and training strategy across all tasks, our method outperforms general-purpose and vessel-specific approaches, with strong preservation of thin branches and vascular continuity. This work not only advances the state-of-the-art in 3D curvilinear segmentation but also opens new avenues for efficient, generalizable, and clinically applicable methods in medical image analysis.
Human lesion studies offer one of the most direct routes to investigating the relations between brain regions and behavioral outcomes in circumstances where experimental interventions are highly restricted. However, these studies face a major challenge in identifying the right level of granularity at which brain regions and behavioral outcomes should be analyzed to identify the relation between lesions to specific brain regions and specific behavioral outcomes. Here we showcase a novel data-driven approach, Causal Feature Learning (CFL), that learns the appropriate level of analysis and the relation between lesion and cognitive impairment at the same time. The method avoids specifying brain regions and specific outcome measures a priori, allowing for the discovery of new cross-cutting lesion-behavior maps. We show that CFL robustly recovers lesion behavior maps in a simulated dataset where Canonical Correlation Analysis fails to provide interpretable results. We then show that CFL recovers known lesion-behavior maps for language deficits and visuospatial processing using a large dataset of lesion subjects, and we illustrate how CFL can be used to identify new groupings of outcomes when mapping lesions to depression symptoms.
Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity. However, whether entropy-based measures reflect clinically meaningful ambiguity, i.e. case-level disagreement about whether a pathology is present at all, remains poorly understood. Contrary to most prior work, which focused on pixel-wise boundary disagreement, we systematically evaluate how well aleatoric uncertainty captures presence ambiguity. Our evaluation spans 3D lung nodule segmentation across four architectures with Monte Carlo dropout and deep ensembles, on LIDC-IDRI and an external validation cohort (LNDb). We find that entropy-based uncertainty maps align with boundary noise and minor drawing variation but carry insufficient discriminative signal for presence ambiguity. In contrast, a lightweight supervised ambiguity head trained on frozen segmentation features substantially outperforms all entropy-aggregation-based baselines across architectures, metrics, and both cohorts, and matches or exceeds methods that explicitly model ambiguity under disagreement supervision (Probabilistic U-Net, Annotator-Confusion 3D-UNet). A qualitative feature-space analysis shows that presence ambiguity is already encoded in the frozen encoder features of pixel-wise-trained networks, only to be discarded by the segmentation output and its entropy aggregation. Our findings expose a fundamental mismatch between the theoretical promise of aleatoric uncertainty and its practical behavior, and suggest that practitioners should not rely on entropy-based uncertainty as a proxy for clinical ambiguity in safety-critical applications.
Particulate matter exposure impacts asthma and chronic obstructive pulmonary disease (COPD) outcomes, and this risk is expected to grow. Therapies that target disease and exposure risk mechanisms are necessary to address this concern. We identified airway epithelial transcriptional mechanisms associated with exposure and disease risk using a functional genomics pipeline that combined unbiased nascent RNA sequencing, genetic prioritization, and gene-regulatory element correlation. We then applied functional studies to prioritized regulatory elements and transcriptional mechanisms and established hyaluronic acid metabolism and TOMM7 regulation as plausible targets supported by airway models and primary human data. Our pipeline shifts the objectives of exposure and genetic studies from broad discovery to refining specific targets, which are more focused on therapeutic manipulation. While future studies are necessary to validate the targets, our framework integrates exposures, transcriptional biology, and human genetics to improve prioritization and streamline translation between target discovery and validation.
Background. Imaging transcriptomics routinely asks whether a trait-associated gene set is over-expressed in a region of interest, and the field standard is to guard that inference with a spatial-autocorrelation-preserving spin test. Electroencephalographic (EEG) oscillatory power is among the most heritable human neurophysiological traits, and the cortical generators of the alpha rhythm have been characterised independently from resting-state magnetoencephalography - making this a natural test bed both for asking whether trait genetics is regionally organised, and for asking what such a test actually establishes. Objective. To test whether alpha-associated genetic signal is spatially enriched in the cortical generators of the alpha rhythm, and to evaluate that inference against complementary null models. Methods. MAGMA gene-based analysis of ENIGMA-EEG summary statistics for six phenotypes: central and occipital alpha power, occipital alpha peak frequency, and theta, beta and delta power. Regional transcription was obtained from the Allen Human Brain Atlas (AHBA) with abagen in the Glasser HCP-MMP1.0 atlas, with Schaefer-100 and Yan-600 as sensitivity analyses. Enrichment in the 41 cortical alpha-source regions was quantified with a threshold-free continuous score and a top-100 gene-set composite, and assessed against three complementary nulls: a spin test (10,000 rotations, cross-checked against brainsmash surrogates), a co-expression-aware gene-set null (10,000 matched random gene sets), and a positive control on a known expression gradient. Results. Judged by the field-standard spin test alone, this study would have reported a positive, biologically coherent finding: alpha-power genes are enriched in the cortical alpha generators (continuous p_spin = 0.022; top-100 p_spin = 0.030), with the spin result corroborated by an independent surrogate model (p = 0.018) and the pipeline validated by a positive control (p_spin = 2 x 10^-4). Three further tests dissolve that conclusion. First, it is not band-specific: theta, beta and delta enrich comparably or more strongly (top-100 beta p_spin = 0.011; delta 0.042; continuous theta 0.042). Second, it does not replicate across alpha phenotypes (occipital alpha continuous p_spin = 0.20; alpha peak frequency non-significant, p_spin >= 0.066). Third, against random gene sets of matched size the alpha set is unremarkable (p_geneset = 0.33) - the apparent enrichment is a generic property of arbitrary gene sets in this cortical territory, and is invisible to a spatial-only null. No test survived false-discovery-rate correction across the 24-cell phenotype x score x region-set grid (minimum q = 0.127), and nominal significance did not survive a change of parcellation. Conclusion. EEG alpha-power genetics shows no regionally specific transcriptomic signature in the cortical generators of the rhythm; the weak tendency that is present is shared across frequency bands, consistent with their known genetic correlation. Methodologically, this is a worked demonstration that correcting for spatial autocorrelation is necessary but not sufficient: a spin-significant, surrogate-corroborated, mechanistically plausible enrichment can be fully accounted for by gene-set co-expression. Enrichment claims in imaging transcriptomics should report a gene-set null alongside the spatial null.
The gold standard datasets for mapping connectomes are electron microscopy volumes of densely labeled neural tissue at nanometer resolution. Yet reconstructing and proofreading neuronal arbors and annotating all synapses requires pipelining multiple software tools that are often fragmented, inconsistently maintained, or proprietary, hindering reproducibility and automation. Here, we introduce Catena, an open-source, comprehensive, developer-centric software suite for connectomics that integrates modules for 3D neuron and organelle segmentation, synapse detection, microtubule tracking, and neurotransmitter inference. Catena organizes its modules in composable, chunk-wise processing pipelines in a completely documented, extensible, and adaptable design. We further reduce compute and ground-truth data requirements with pretrained machine learning models, facilitating fine-tuning. Catena ships fully containerized modules that encapsulate evolving dependencies for consistent execution across workstations and clusters. By consolidating open components, shareable models, and containerized runtimes, Catena delivers a reproducible and scalable approach to mapping cellular connectomes from electron microscopy volumes. Code and documentation: https://github.com/Mohinta2892/catena.git
Brain lesion segmentation is a fundamental task in medical image analysis, playing a critical role in diagnosis, treatment planning, and longitudinal disease monitoring. Yet existing models still struggle to meet the demands of real clinical use, where deployments contain data distribution shifts, arising from differences in scanner hardware, imaging protocol (varying MRI modality sets), and new pathologies. We present BrainIAC (Brain lesion Interactive Adaptive Continuously learning segmentation), a unified framework that integrates (i) a multi-modal backbone network trained to segment multiple types of brain lesions and handle heterogeneous sets of modalities via zero-filling and random modality dropping; (ii) 3D interactive segmentation with bounding-box and click prompts that preserves fully automatic prediction when no prompt is given; and (iii) an online adaptation mechanism combining Mid-Interaction adaptation and Post-Interaction adaptation, supervised by the network's own predictions as pseudo labels and guided by an extra Click-Centered Gaussian loss. To our knowledge, this is one of the first 3D online adaptation methods for interactive segmentation, and the first to combine handling of heterogeneous modality sets with online adaptation. Experiments across seven brain MRI datasets demonstrate that the proposed components provide complementary and synergistic benefits. The method consistently outperforms existing approaches and generalizes well across heterogeneous imaging modalities, including those unseen during training, as well as previously unseen brain pathology types. The code and a 3D Slicer plug-in will be released at https://github.com/WenTXuL/BrainIAC upon publication.
Purpose: Contrast-enhanced breast MRI depends on intravenous gadolinium-based contrast agents, motivating methods that synthesise post-contrast appearance from pre-contrast images alone. The MAMA-SYNTH Challenge (MICCAI 2026 Deep-Breath Workshop) evaluates such synthesis across four metric groups: image fidelity, tumour region of interest, downstream classification, and downstream segmentation. The four group ranks are averaged, so optimising a single objective is insufficient. Materials and Methods: We developed a lesion-gated hybrid synthesis pipeline. A lesion probability map, estimated from the pre-contrast slice alone by an ensemble of pre-contrast-only segmentation networks, spatially coordinates a tumour-focused regression pathway and a background-focused Pix2PixHD synthesis pathway, followed by a region-dependent calibration of the predicted enhancement. Training used the public MAMA-MIA collection with a patient-level split. Inference consumes a single pre-contrast 2D slice, with no mask, no post-contrast image, and no auxiliary metadata beyond image geometry. Results: On internal validation (n = 120 patients) the method reached MSE 0.523, LPIPS 0.185, and tumour-region SSIM 0.495. It was submitted to the hidden external 300-case test cohort as a self-contained inference container. The method ranked sixth in the official MAMA-SYNTH Challenge leaderboard. Conclusion: A test-compatible lesion probability map derived from the pre-contrast image can coordinate complementary tumour-focused regression and background-focused perceptual synthesis, enabling balanced virtual contrast enhancement under a multi-metric challenge setting.
Despite advances in understanding the mechanistic underpinnings of chronic pain (CP) conditions, knowledge on the brain anatomy microstructure associated to CP trajectories is still scarce. We sought to fill this knowledge gap by combining longitudinal self-reported CP assessments with brain imaging in a community dwelling adult population (n = 976). Aiming to provide anatomical insights beyond morphometry (based on T1 brain images), we analyzed the derived relaxometry measures MTsat, R1 and R2* and we could investigate myelin and iron content. Participants were divided in three groups according to persistent, resolved or new CP and were compared to those without CP at timepoint. Participants with persistent CP exhibited lower grey matter (GM) in fronto-temporal cortical regions and reduced myelin-related signal in subcortical and thalamic areas, alongside higher iron-sensitive metrics in the entorhinal cortex. Those who recovered from CP showed reduced cortical GM and myelin encompassing subcortical nuclei, thalamus, frontal, parietal, and temporal regions. New-onset CP participants displayed reduced GM in cortical areas. Overall, we observed reduced GM volume in a widespread network of fronto-temporal cortical areas in participants with current, resolved and new-onset CP. Myelin-related metrics were lower in the subcortical nuclei and thalamus in participants with persistent and -to a greater extent- with resolved CP, whilst persistent CP was associated with higher iron-sensitive metrics in the entorhinal cortex. Hence, CP trajectories appear to leave differential imprints on brain morphometry and tissue microstructure, with each trajectory associated with distinct non-overlapping spatial and temporal patterns in GM, myelin, and iron content.
Counterfactual image generation answers questions about how a subject would have looked under retrospective, hypothetical scenarios. Recent methods have improved perceptual quality, identity preservation and faithfulness to an underlying causal model, but their adoption in healthcare is limited by scarce annotated data, distribution shift between datasets, and mismatches between pretrained generative models and those required for counterfactual inference. We propose specialisation, a data and parameter-efficient framework for adapting pretrained, non-causal generative models into causal mechanisms under distribution shift. Based on this framework, we train a radiology counterfactual image generation model, called RadCF, using latent flow matching. We validate our approach on three chest X-ray datasets spanning different dataset shifts, data volumes, and counterfactual questions, associated with challenging, highly-localised interventions. Our results show that RadCF and specialisation improve counterfactual soundness over existing methods while being data and parameter efficient, and that the resulting counterfactuals can detect and mitigate shortcut learning in a downstream medical classifier. Code is available at https://github.com/GSK-AI/RadCF/.
Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the latent codes for image generation to operate on. Thereby, the tokenizer choice bounds every downstream task from reconstruction fidelity and generation quality to the representations available for downstream analysis. Yet, medical imaging pipelines routinely utilize tokenizers from natural imaging on the hypothesis that their behavior carries over. However, this is an assumption never tested in the medical imaging regime, where datasets are orders of magnitude smaller and images exhibit far lower inter-sample variance. We present a systematic evaluation of medical image tokenizers evaluating thirty configurations across ten model families on twelve datasets at three compression factors, spanning reconstruction, generation, latent geometry, downstream classification, and memorization. We find that (1) performance on image reconstruction and generation strongly correlate, unlike prior reports on natural images; (2) modern tokenizers use nearly all of their codebook entries, but still leave most of the latent space unused; (3) training-set memorization is mild and is further suppressed by stronger latent space compression; and (4) discrete quantization can largely preserve downstream classification, with lookup-free schemes being the main exception.
🏫 主要单位:Virginia Institute for Psychiatric and Behavioral Genetics, Department of Psychiatry, Virginia Commonwealth University, Richmond, VA, USA | 主要作者:Tan-Hoang Nguygen、Kramer, S.
Post-traumatic stress disorder (PTSD) has a significant genetic component (h2 = 24-72%). While the majority of prior studies have utilized common variants, recent research suggests rare variants also contribute to PTSD liability. Here, we examined the role of rare copy number variants (rCNVs; MAF < 1%) from multi-ancestry whole-genome sequencing data in the All of Us Research Program, using an electronic health record-defined PTSD phenotype. Genome-wide burden analyses utilizing all rCNV lengths revealed significant associations for deletion count across European-like (EUR-like; beta = 0.005, SE = 0.001, P = 5.99 x 10-6, FDR P = 3.59 x 10-5 ) and African-like (AFR-like; beta = 0.003, SE = 0.001, P = 0.005, FDR P =0.02) ancestries. We conducted rCNV burden analyses within and across EUR-like (N case = 1,179, N control = 33,689), AFR-like (N case = 504, N control = 12,216), and Admixed American (AMR-like) (N case = 418, N control = 11,935) ancestries using three separate domains of gene-sets previously implicated in neuropsychiatric disorders: neurodevelopmental disorders (NDD; N = 53), abnormal behavior mouse mutant-derived (N = 146), and PTSD (N = 5). Within-ancestry analyses across all three domains yielded one significant brain-expressed gene-set within the EUR-like cohort. Meta-analyzing across ancestries identified 11 NDD gene-sets enriched for rCNV count (FDR P < 0.05). rCNVs with lengths greater than 10kb were examined, but observed no significant signals. Future research using larger samples and detailed PTSD phenotypes, integrated with functional genomic data and brain-region-specific expression profiles, is needed to fully map the disorder's genetic landscape.
Acquiring high quality annotated medical image data is critical for training deep learning models; however, annotation is expensive, time consuming, and requires domain expertise. Conditional diffusion models, such as ControlNet, offer an alternative by generating images conditioned on semantic masks and text. However, existing approaches fail to capture fine grained properties (e.g., intensity and texture), as well as semantic consistency expected by domain experts, limiting their effectiveness for downstream tasks. Recent attempts to address these issues using reinforcement learning fine-tuning remain limited due to the reliance on a single scalar reward, which conflates diverse failure modes and provides weak corrective signals. We propose PRISM, a Compositional Reward Model (CRM) framework for conditional medical image generation. Instead of assigning a single reward, we decompose image quality into verifier grounded stages, each evaluating a distinct aspect of correctness from fine to coarse properties, including low level attributes (intensity and texture), structural alignment with conditioning inputs, and high level semantic fidelity. These stage wise rewards are composed through a Hierarchical Constrained Propagation (HCP) mechanism that enforces a fine to coarse notion of correctness, ensuring that lower level deficiencies are resolved before higher level rewards are accrued, preventing easier objectives from masking critical failures. We evaluate PRISM across three datasets spanning diverse medical imaging tasks: PanNuke (multi-class cell segmentation), CeDeM (villi/crypt detection and measurement), and ISIC (skin lesion classification). Training downstream models with data generated by PRISM yields improvements over closest baselines, including a 2.3% increase in mDice on PanNuke, a 8.5% reduction in Mean Relative Error (MRE) on CeDeM, and increases ISIC F1 by 5.9%.
While large language models are powerful generators of new text, forecasting disease progression from longitudinal health histories remains a challenging problem. We introduce GenEHR, an autoregressive generative model trained on electronic health records (EHRs) from millions of patients that explicitly represents the irregular time intervals between visits when forecasting future clinical events. We combine the general-purpose patient representation learned during foundational training with parameter-efficient supervised adaptation for the task of pan-cancer risk stratification. In five large EHR cohorts supervised adaptation substantially improved prediction performance of a first cancer diagnosis within a five year horizon window. Our retrospective results support the evaluation of GenEHR-CancerRisk as a prospective clinical decision-support tool for prioritizing patients for risk-based screening for aggressive cancer types, such as pancreatic and ovarian cancer.
Accurately isolating disease-related features from confounding covariates (e.g., age, gender, site) and individual variations remains a fundamental challenge in medical image classification. Traditional regression-based approaches may ignore non-linear relations between image features and true covariates. To overcome this issue, we present a generalized Medical Imaging Disentanglement Learning (MedIDL) framework. MedIDL maps image features into three mutually orthogonal latent spaces through specialized disentanglement heads: a disease classification head guided by a supervised loss, a covariate-alignment head constrained by cross-subject similarity matching, and a Gaussian head absorbing individual variations. We evaluated our framework across 7 datasets encompassing diverse imaging modalities. MedIDL outperforms state-of-the-art supervised and self-supervised classification methods in accuracy across all datasets. Association analyses demonstrate that MedIDL successfully isolates target-specific latent representations. Gradient-based interpretability mappings localize pathognomonic patterns aligning with established clinical literature.
BackgroundHistopathological diagnoses in kidney transplant biopsies, as described by the Banff classification, suffer from substantial inter-observer variability and misinterpretation, making them prone to diagnostic errors. Continuous indices based on BanffNET automated lesion scores are proposed to quantify the continuous phenotypic spectrum of kidney transplant biopsies, complementary to pathologist assessment of biopsies. MethodsThe BanffNET model was used to obtain continuous lesion scoring for a training cohort of 2544 biopsies, a large validation cohort of 3863 biopsies, and a multi-reader cohort of 36 biopsies scored by 67 pathologists. A penalized regression approach was used to condense 17 automated lesion scores into four indices, representing the most important dimensions of graft injury: a TCMR/TI Index, an AMR/MVI Index, an Activity Index and a Chronicity Index. The BanffNET Indices were evaluated against Banff diagnoses and histological indices in the training cohort and validation cohort. Since pathologist-assigned Banff lesions and diagnoses suffer from inter-observer variability, associations were also assessed with diagnoses in a multi-reader study, with time to kidney graft failure in the training and validation cohort and with biopsy-based molecular signatures in the validation cohort. ResultsThe BanffNET TCMR/TI Index, AMR/MVI Index and Activity Index showed excellent discrimination of Banff TCMR, Banff MVI and Banff Any Diagnosis respectively, with validation AUCs ranging from 0.80 to 0.92. In addition, the BanffNET Indices explained large amounts of variability in multi-observer reference standards, time to graft failure and molecular biopsy-based diagnostics. In the large majority of cases, variability explained by BanffNET Indices matched or exceeded the variability explained by pathologist-assigned diagnoses and histological indices. ConclusionsBanffNET Indices are reproducible descriptors of whole slide images that could complement and augment kidney transplant biopsy evaluation by pathologists. Associations of BanffNET Indices with multi-observer diagnoses and related outcomes suggest that they capture additional information on tissue injury type and severity on top of pathologist-assigned Banff diagnoses and histological indices.
Unsupervised anomaly detection (UAD) aims to localize abnormal regions in medical scans without pixel-level annotations. A typical strategy seeks to reconstruct a pseudo-healthy image that preserves subject-specific anatomy. Recently, diffusion models have been proposed to perform UAD. However, these methods rely on heuristic noise schedules or synthetic corruptions to balance subject-specificity and anomaly removal. In this work, we propose an alternative formulation of UAD as a Bayesian inverse problem under a diffusion prior. First, we introduce a latent spatial anomaly mask that models pixel-wise consistency between a test image and its latent corresponding pseudo-healthy image. Then, we propose an approximation of the unknown generation process that links healthy anatomy, anomalies, and the observed image, enabling a well-defined likelihood within the Bayesian framework. Building on recent advances in diffusion-based inverse problem methods, we jointly infer the pseudo-healthy image and the anomaly mask via annealed posterior sampling. We evaluate our approach on FDG PET (ADNI) and FLAIR MRI (BraTS 2021), demonstrating improved anomaly localization performance compared to other diffusion-based approaches and validating the contribution of our introduced model. Our code is available at https://github.com/HuguesRoy/UAD_DAPS.
INTRODUCTION: Determining whether plasma biomarkers are preferentially associated with Alzheimer's disease (AD)-related rather than age-related brain atrophy patterns may clarify their prognostic and diagnostic clinical use. METHODS: Using data from the Baltimore Longitudinal Study of Aging (N=818), we examined cross-sectional plasma A{beta}42/A{beta}40, GFAP, NfL, p-tau181, and p-tau217 measurements obtained while participants were cognitively unimpaired (CU). During follow-up, 104 participants developed mild cognitive impairment (MCI)/dementia (74 due to AD, 24 due to non-AD, 6 unknown etiology). 2,293 longitudinal brain MRIs were used to quantify multidimensional atrophy pattern scores reflecting brain age (SPARE-BA), AD-like patterns (SPARE-AD), and five dominant dimensions of atrophy (R-indices). We investigated the associations of plasma biomarkers and atrophy pattern scores at index visit with conversion to MCI/dementia due to AD. We then examined the associations of plasma biomarkers with longitudinal change in pattern scores using linear mixed effects models. RESULTS: p-tau181, A{beta}42/A{beta}40 (Lumipulse), and p-tau217 were associated with incident MCI/dementia due to AD but not non-AD etiologies, while GFAP was associated with incident MCI/dementia due to both AD and non-AD. All four biomarkers were associated with longitudinal SPARE-AD changes. A{beta}42/A{beta}40 (Quanterix and Lumipulse), p-tau181, and p-tau217 were associated with longitudinal parieto-temporal atrophy. p-tau217 was the only biomarker associated with longitudinal medial temporal lobe atrophy. We did not find associations between plasma biomarkers and SPARE-BA or R-indices capturing subcortical, diffuse cortical, or perisylvian atrophy. DISCUSSION: Among CU individuals, plasma p-tau217 was associated with subsequent MCI/dementia due to AD and showed the most extensive associations with longitudinal AD-related brain atrophy.
Accurate classification of brain disorders from neuroimaging data remains challenging because of substantial inter-subject heterogeneity and the complex multi-scale patterns present in functional connectivity and morphological representations. To address these challenges, we propose HyperAMS-Net, a deep learning framework for brain disorder classification using neuroimaging representations derived from resting-state functional MRI or structural MRI. HyperAMS-Net integrates adaptive multi-scale convolution, hypergraph attention, spatial-channel attention, and adaptive feature fusion. Specifically, adaptive multi-scale convolution learns data-driven weights over multiple receptive fields to capture complementary patterns at different scales. Hypergraph attention models higher-order dependencies among learned feature representations through node--hyperedge--node message passing, while spatial-channel attention enhances discriminative feature learning. Adaptive feature fusion further aggregates complementary information across parallel network branches. HyperAMS-Net is evaluated on three benchmark datasets spanning distinct brain disorders: ABIDE for autism spectrum disorder, REST-meta-MDD for major depressive disorder, and ADNI for Alzheimer's disease, using 5-fold stratified cross-validation. HyperAMS-Net achieves state-of-the-art performance across all evaluated datasets, attaining the highest accuracy and AUC among the compared methods. Ablation studies further demonstrate the contribution of each proposed component, with the largest performance degradation observed when hypergraph attention is removed.
Recent innovations in deep learning have significantly enhanced the diagnosis of medical images, although they are based on the use of centralized data storage that pose severe threats to patient privacy and medical data security. To address this issue, this research proposes a Federated Learning (FL) model that is coupled with an optimized YOLOv8 network to detect the kidney stones on a computed tomography (CT) image and at the same time, protect privacy of the patients. The suggested system can help various medical organizations to jointly train a common model without exchanging the information about the patients. This is to ensure that data protection laws like GDPR and HIPAA are adhered to. The residual feature fusion and DropBlock regularization among other architectural improvements are also included in YOLOv8 to enhance detection robustness and minimize overfitting. Experimental analysis carried out on a distributed CT dataset demonstrated that the federated YOLOv8 model has a mAP at 50 of 0.733 and is able to keep the data confidential. Moreover, its lean design facilitates fast edge deployment and real-time inference across a clinical setting. Altogether, these findings indicate that Federated Learning is a safe and efficient solution to AI-assisted diagnosis in contemporary healthcare when combined with the use of sophisticated object detection models.
AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (MoE) architectures currently fail to satisfy. Expert-based routing improves in-domain learning but may sacrifice cross-modal signals, which appear particularly important for rare (low-prevalence) pathologies in our experiments. To resolve this, we introduce Generalist-Specialist-MoE (GS-MoE), a two-branch (MoE) architecture that couples a cross-modal generalist model with distinct modality-specific specialists (experts) via domain-constrained feature fusion. On RadImageNet (1.35M images, 165 pathologies, three modalities), GS-MoE recovers detection of six low-prevalence pathologies on which every baseline scores F1 $=$ 0, with per-class gains up to +0.60 F1. It attains this while even slightly exceeding dense and specialist-only MoE aggregate baselines (MCC 0.770), while using ${\sim}53\%$ fewer active parameters at inference than the strongest investigated dense model.
Donor-derived cell-free DNA (dd-cfDNA) is increasingly used for post transplantation non- invasive surveillance; however, its clinical interpretation remains inconsistent, with widely ranging thresholds and is typically applied as a single binary cutoff in literature. The optimal decision framework for rule-out and rule-in decisions, and whether a single threshold remains clinically meaningful, are currently uncertain. We performed a Bayesian hierarchical summary receiver operating characteristic (HSROC) meta-analysis of 14 studies (1,763 patients) evaluating dd-cfDNA against endomyocardial biopsy. To account for serial testing within individuals, we applied a cluster-corrected design effect, reducing 6,103 observations to 2,518 effective tests. Threshold-dependent sensitivity and specificity were modelled continuously. We compared a conventional single-threshold approach with a data-driven adaptive framework defining rule-out and rule-in thresholds and evaluated clinical utility by decision-curve analysis across rejection prevalences from 1% to 50%, incorporating repeat-testing strategies. The pooled area under the HSROC curve was 0.78 (95% CrI, 0.67-0.84). The Youden-optimal threshold (0.20%) yielded balanced sensitivity (0.77) and specificity (0.77) but failed to support clinical objectives of diagnosis. An adaptive framework identified a rule-out threshold of 0.16% (sensitivity 0.80) and a rule-in threshold of 0.48% (specificity 0.90), defining a indeterminate / grey zone. The residual one-in-five false-negative rate at the rule-out anchor reflects low-grade, non-cytolytic rejection, the imperfect histological reference standard and fractional suppression of the donor signal, rather than the statistical model; a result above the rule-in anchor carries a positive predictive value of approximately 38% at 10% prevalence and denotes an indication for tissue diagnosis and multimodal investigation, not for empiric treatment. Across low-to-intermediate prevalence, dd-cfDNA-guided strategies exceeded both the biopsy-all and monitor-all reference strategies; among testing strategies, repeat-if-borderline achieved the highest net benefit across the majority of the prevalence-threshold space and sustained positive net benefit over the widest operating range of any strategy, reducing false-positive biopsies without materially compromising detection. At high prevalence, where a first elevated result is usually true, biopsy-all became competitive. A single threshold is therefore clinically inadequate for post-transplant surveillance. Our tri-state, prevalence-aware framework integrating rule-out, indeterminate, and rule-in zones with selective repeat testing, more accurately reflects biomarker behavior and yields greater net benefit than any single cutoff because these anchors are pooled, population-level estimates rather than universal constants, programs should adopt this architecture and calibrate their own high-sensitivity rule-out and high-specificity rule-in thresholds to their local assay and population.
A medical model's benchmark score does not establish that the same conclusion holds under a different evaluation. This study tests whether claims about model ranking, score reliability and screening performance survive changes in cohort, prompt, negative spectrum, specified prevalence and operating threshold. We audit three medical vision-language models (BioMedCLIP, CheXficient, and MedSigLIP) and a general-domain OpenCLIP comparator on 12,200 chest radiograph records from four datasets (Montgomery, Shenzhen, TBX11K, and VinDr-CXR). Five fixed prompt families yield 244,000 model--image--prompt scores. No model leads every cohort and reliability criterion. Prompt-family changes alter AUROC in 21 of 48 multiplicity-controlled comparisons. Replacing healthy controls with sick non-tuberculosis controls reduces AUROC by 0.075--0.306 across all four models. On VinDr-CXR, the three medical models distinguish tuberculosis from no-finding controls substantially better than from pneumonia or lung tumor; their AUROC point estimates for both named diseases fall below 0.5. CheXficient has documented VinDr-CXR pretraining exposure, which limits the interpretation of its results. Thresholds chosen for 95\% sensitivity on TBX11K training retain that constraint by point estimate in only four of sixteen target evaluations. A five-seed supervised source model reaches 0.999 AUROC on TBX11K validation but 0.629 on each of two external cohorts. Conservative exclusion of perceptual-overlap candidates narrows this gap without closing it. These retrospective, single-task results show that discrimination, score reliability and threshold retention support different portability claims. Evidence for chest X-ray tuberculosis screening should identify the complete evaluation specification rather than attribute clinical portability to a checkpoint alone.
Multi-label few-shot learning (MLFSL) remains a significant challenge in medical image analysis (MIA). Current metric-based meta-learning methods face two critical limitations in MIA. First, conventional prototype generation often entangles irrelevant disease information, leading to contaminated prototypes and degraded performance. Second, prior studies typically enforce inter-class separability in embedding space, largely neglecting the inherent correlations among diseases. To overcome these challenges, we propose Prototype Purification and Regulation (PPR), a novel MLFSL framework for MIA. PPR first performs prototype purification by leveraging sample-level comorbidity scores to emphasize disease-specific features, producing purified prototypes that better characterize each disease. Building upon these purified prototypes, PPR further addresses the underexplored problem of inter-class prototype distance in MIA by incorporating disease-level comorbidity statistics to adaptively regulate inter-class similarity, forming a comorbidity-aware embedding space. Overall, PPR sequentially enables the model to capture pure disease features and inter-class relationships for reliable MLFSL in MIA. Extensive experiments across four chest X-ray benchmark datasets, including cross-domain evaluation, show that PPR consistently outperforms state-of-the-art methods, significantly improving disease detection while demonstrating robust generalization and clinical applicability.
In this work, we present an uncertainty-driven training framework for three-dimensional computed tomography (CT) lung nodule classification, where validation-based uncertainty estimates guide loss reweighting to enhance predictive performance and probability calibration. Two Uncertainty Quantification (UQ) methods are considered: Monte Carlo Dropout (MCD) and Evidential Deep Learning (EDL). Both provide per-class uncertainty estimates that modulate the loss and encourage focus on hard or unreliable classes. The framework is evaluated with ResNet, DenseNet, EfficientNet, Vision Transformer (ViT), and Swin Transformer backbones on two datasets: the clinical LIDC-IDRI cohort and the NoduleMNIST3D benchmark. Uncertainty-driven training achieves classification performance similar to conventional training while substantially improving calibration, with an expected calibration error (ECE) reduced by up to 65% on LIDC-IDRI. EDL attains competitive performance on shallower architectures with single-pass inference, whereas MCD is more robust on deeper networks. Analysis across architectural families reveals that uncertainty-driven training benefits convolutional backbones more consistently than transformer-based architectures: EDL in particular degrades on ViT, suggesting that the Dirichlet evidence parameterisation may interact unfavourably with attention-based architectures at lower input resolutions. A posteriori temperature scaling proves highly effective across all configurations, indicating that a simple scalar calibration can be competitive even without explicit uncertainty-aware training. Our results indicate that integrating UQ into the training loop can significantly improve probabilistic calibration and support more trustworthy deployment of three-dimensional medical imaging models.
Background: Cascade screening increasingly identifies carriers of pathogenic or likely pathogenic sarcomere variants at risk for hypertrophic cardiomyopathy (HCM) in whom penetrance is incomplete, and surveillance relies on resource-intensive serial imaging. We evaluated whether a validated artificial intelligence-enhanced electrocardiography (AI-ECG) model identifies the HCM phenotype at first clinical assessment, predicts development of HCM during follow-up, and complements polygenic risk. Methods: We assembled 1,095 genotype-positive (G+) individuals with pathogenic or likely pathogenic sarcomere variants from Yale-New Haven Hospital (n=119), Erasmus MC (n=858), and Motol University Hospital (n=118). At baseline (first clinical assessment), individuals were classified as phenotype-positive (P+) or phenotype-negative (P-). A previously validated AI-ECG model applied to 12-lead ECG images generated an HCM score. The primary outcome was detection of phenotypic positivity at baseline; secondary analyses included manifest HCM (at baseline or during follow-up) and identifying risk of developing future HCM among G+/P- individuals. In 57,007 UK Biobank participants, we assessed whether AI-ECG adds to an established polygenic risk score (PRS). Results: Among 1,095 G+ individuals (median age 46 years [IQR 34- 56]; 52.1% female), 808 (73.8%) were P+ at baseline, 56 (5.1%) developed HCM during follow-up, and 231 (21.1%) remained P-. AI-ECG achieved an AUROC of 0.91 (95% CI 0.89-0.93) for P+ at baseline and 0.92 (95% CI 0.90-0.94) for manifest HCM. At a threshold of 0.15, sensitivity was 0.78, specificity 0.89, PPV 0.95, and NPV 0.59. Among G+/P- individuals, higher AI-ECG scores predicted development of HCM (HR 1.55 per 1-SD; 95% CI 1.28- 1.88; p< 0.001; adjusted HR 1.38; 95% CI 1.11- 1.71; p=0.004). In the UK Biobank, individuals with both high AI-ECG and high PRS had 60-fold higher odds of HCM (adjusted OR 60.2; 95% CI 26.5- 137.2), versus 15.0 for high AI-ECG alone and 4.1 for high PRS alone. Conclusions: AI-ECG detects the HCM phenotype at baseline in sarcomere variant carriers, predicts development of HCM in G+/P- individuals, and complements PRS in the general population, supporting AI-ECG as a scalable tool to detect HCM and guide surveillance in individuals with monogenic or polygenic susceptibility.
Developing competitive deep learning baselines for medical imaging remains a highly iterative process requiring literature review, implementation, experimentation, and expert refinement. Existing automation approaches typically optimize isolated components, such as architecture search or hyperparameter tuning, rather than the complete baseline development process. We present an agentic AI Scientist workflow that combines literature-guided reasoning, automated code generation, and hypothesis-driven experimentation to generate competitive baseline models for medical imaging challenges. The framework is evaluated on four public benchmarks spanning segmentation, classification, and detection. Across all tasks, the Experimentation Pipeline consistently improves validation performance, achieving competitive leaderboard results, including 6th place on both PUMA tracks (15 teams) and 31st place on MILK10k (125 teams). On MIDOG25, the resulting model also demonstrates strong domain generalization across scanners, tumor types, and species. Using the same workflow across all challenges without task-specific redesign, we demonstrate that skill-based, literature-guided agentic workflows can substantially reduce the engineering effort required to develop competitive medical imaging baselines.
Undergraduate programmes in intelligent medical engineering are expanding, yet laboratory curricula lag behind the multimodal, long-tailed, and distributionally shifting realities of clinical AI. This design paper presents an advanced experimental teaching system that translates an ongoing multimodal deep learning research project on endometrial carcinoma into a structured undergraduate lab sequence. We identify three educational gaps (modality, authenticity, and deployment) and derive four pedagogical principles from constructive alignment, experiential learning, the research teaching nexus, and the CDIO framework. The curriculum comprises four progressive tiers plus an engineering layer, with 32 laboratory units over 64 contact hours, delivered via a custom virtual clinical workstation using de-identified multi-institutional data. Each tier maps to a specific technical bottleneck, prerequisite coursework, and criterion-referenced deliverables. Data governance, safety, and assessment protocols are specified. Learning outcome data will be collected across two implementation cycles.
Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local service retains authority to accept or replace that proposal under referral policy, capacity, and locally held patient context. Proposal quality and selected-action quality are therefore distinct evaluation targets, and benchmark evidence transfers between them only when local review preserves the expected action score. We introduce AuthEval, a logging and evaluation framework that records both actions, scores the selected action under declared local criteria, and, where feasible, scores the declined proposal under the same rule. It reports the resulting authority gap only when the record supports it. Because the gap is the product of the proposal-change rate and the mean score change on changed cases, that rate alone determines neither its magnitude nor its sign. On ISIC 2019, with MedGemma and simulated local review, two constraint regimes with similar change rates produced an optimistic image-equal gap under capacity ($+0.744$ simulator units) but no detectable gap under safety. The declared evaluation unit also mattered: under mixed constraints the gap reversed from $+0.374$ to $-0.206$ when weighting shifted from image to lesion-aware cluster. AuthEval thus clarifies whether a study's records support claims about the model, the workflow, or both.
This work proposes a real-time 6 Degree-of-Freedom (6 DoF) tracking algorithm for monocular laparoscopic surgery. It provides alignment between intra-operative video and pre-operative data (e.g., CT). The 6 DoF tracking offers a solution for accurately locating the internal anatomy of the target organ despite the lack of tactile feedback and transparency. The ORB-SLAM2 framework is adopted and modified for prior-based 3D tracking with four major modifications. First, the primitive 3D shape is used for fast initialization of the ORB-SLAM2 monocular mode. Second, a pseudo-segmentation strategy is employed to separate the target organ from the background for tracking. Third, the 3D shape is incorporated as a geometric prior in its pose graph optimization. Fourth, the Multi-Scale Retinex with Chromaticity Preservation (MSRCP) algorithm is leveraged and modified for image enhancement in challenging illumination scenarios. In-vivo and ex-vivo experiments validate that LapaTrack-3D provides robust 3D tracking and effectively handles typical challenges such as poor illumination, fast motion, out-of-field-of-view scenarios, partial visibility, and ``organ-background'' relative motion. LapaTrack-3D achieves a processing rate of 13 Hz for 1280*720 pixel video.
We propose a new probabilistic model for general-purpose medical image registration that builds upon the mutual information registration criterion. It centers around a spatial interpolation technique that assumes latent voxel-wise correspondences between the images being registered. By exploiting these latent variables, we derive dedicated optimization and MCMC sampling techniques that only involve closed-form iterative updates. When applied to nonlinear registration, an efficient demons-like optimization algorithm is obtained that shows robust out-of-the-box performance across a variety of monomodal and multimodal registration tasks. We also demonstrate a corresponding sampler that can quantify, for the first time, uncertainty in multimodal registration scenarios with very high-dimensional 3D deformations. Our code, which we call BINDER (Bayesian INference for DEformable Registration), is freely available at https://github.com/ste93ste/BINDER.
We consider a nonlinear model motivated by polychromatic computed tomography (CT). Here, $W$ distinct $d$-dimensional signals $x^*_1, \ldots, x^*_W \in \mathbb{R}^d$ must be recovered from $n$ measurements $(a_i, y_i)_{i = 1}^n$ that obey the nonlinear model $\mathbb{E}[y_i|a_i] = h(\langle a_i, x^*_1 \rangle, \ldots, \langle a_i, x^*_W \rangle)$, where $h: \mathbb{R}^W \to \mathbb{R}$ is a known nonlinearity that models a certain type of exponential attenuation law. Even when there is no noise in the measurements, the sample size $n$ (for any measurement ensemble $\{a_i\}_{i = 1}^n$) must exceed the number of unknowns $Wd$ to guarantee that the underlying signals are identifiable. We construct a measurement ensemble that ensures perfect signal recovery almost surely provided $n \geq Wd + W - 1$, thereby isolating the injectivity threshold up to an additive factor $W - 1$. We also study computational complexity of signal recovery in this model with a general measurement ensemble $\{a_i\}$. We show that if $W \geq 2$, then there is a measurement ensemble for which deciding whether there exist signals consistent with the measurements is NP-hard. In particular, this implies that polychromatic, multimaterial CT reconstruction is NP-hard in general. This finding stands in sharp contrast to the single-material setting, for which a polynomial-time algorithm can provably perform signal recovery for any measurement ensemble.