Transferring the rich priors of large 2D foundation models to sparse 3D LiDAR remains challenging, as training native 3D foundation models at comparable scale is limited by data and annotation scarcity. We introduce a LiDAR-conditioned diffusion model trained on pseudo-labels from off-the-shelf 2D foundation models. The model supports multiple output modalities, including depth, semantic segmentation and instance prediction, selectable via a textual task prompt. Because the model is conditioned on LiDAR, both its outputs and its intermediate UNet features can be projected back onto the input point cloud, enabling analysis of a 3D representation learned entirely under 2D supervision. We study this representation directly in point-cloud space, explicitly excluding raw spatial coordinates to isolate feature content from projection geometry. Linear probes recover up to ~23% Mean Intersection over Union (MIoU) on 3D semantic classes, compared to ~3.5% for a matched Gaussian-noise control, indicating substantial non-trivial structure. Pairwise cosine similarity across modality-specific feature streams reveals a layered organization. Early encoder layers remain weakly aligned across modalities while individually decodable, intermediate layers converge toward a shared representation, and decoder layers re-specialize toward task-specific outputs. These findings indicate that LiDAR-conditioned diffusion models can induce structured 3D representations from 2D supervision alone, with a modality-dependent manifold that locally unifies near a shared bottleneck. This positions diffusion as a viable mechanism for transferring large-scale 2D priors into sparse 3D domains.
🏫 主要单位:Department of Biomedical Informatics and Data Science, Heersink School of Medicine, University of Alabama at Birmingham, AL, USA | 主要作者:Lana X Garmire、Liu, T.
Spatial omics enables the integration of high-dimensional molecular organization with clinical outcomes, yet incorporating spatial single-cell information into predictive models at the population scale remains challenging. Here, we introduce SPIN (Spatial Predictive Integration Network), which integrates subject-specific, spatially informed co-expression or co-abundance networks for population-scale clinical prediction. Current implementation of SPIN includes three key modules: network construction, molecular-to-clinical prediction by Bayesian scalar-on-network regression with manifold learning (BSNMani), and visualization & interpretation module. In the SEA-AD MERFISH transcriptomics cohort, SPIN framework achieved an accuracy of 0.81 for dementia-status prediction and revealed gene-gene co-expression subnetworks with clear biological relevance, such as glutamatergic synapses- and neurogenesis-related subnetworks. SPIN also achieved robust survival prediction in a breast cancer single proteomics cohort (C-index=0.78), identifying two survival-associated spatial proteomic subnetworks. SPIN can further use cell-type-specific spatial omics data to enhance the granularity. In summary, SPIN is a valuable tool that uses high-dimensional spatial omics data for clinical outcome prediction at the population scale across diverse disease settings, revealing biological insights while maintaining interpretation. SPIN package is available at: https://github.com/lanagarmire/SPIN
Pancreatic tumor segmentation in 3D CT volumes is challenged by extreme scale variability across both the pancreas and tumor, and highly irregular tumor morphology. While recent advances have pushed segmentation performance, existing methods do not explicitly address these challenges and come at the cost of excessive computational complexity, limiting their practicality in resource-constrained clinical environments. We propose TRIUNE-Net, a lightweight unified architecture that harmonizes scale, shape, and efficiency through three synergistic innovations. A multi-scale context aggregation module with stage-adaptive dilated convolutions enables the model to reason across the broad range of anatomical scales present in both organs. A serial linear-deformable attention mechanism combines large effective receptive fields with shapeadaptive deformable convolutions to capture irregular, non-convex tumor morphologies. Finally, an information-preserving downsampling module replaces conventional max pooling entirely, retaining all spatial information while adding negligible parameters, preventing small tumors from being discarded before they can be recognized. On both the MSD Pancreas and NVD Pancreas datasets, TRIUNE-Net achieves state-of-theart results with only 5.86 M parameters and no external pre-training, outperforming all baselines across all key tumor metrics. Specifically, it surpasses the next-best model by 0.45% in tumor Dice, 6.0 points in F1 score, 6.6 points in sensitivity, and 3.4 points in precision, simultaneously reflecting its ability to suppress both missed tumors and false alarms in clinically realistic conditions. Our code is available at: https://github.com/abdora-ai/TRIUNE-Net
Accurate 3D plant organ segmentation is fundamental to automated phenotyping. Existing approaches rely on annotated training data or species-specific model configurations. We present an annotation-free pipeline for 3D plant organ segmentation, combining text-prompted SAM3 segmentation with semantic neural radiance fields (NeRFs). Given only multi-view RGB images and a list of class names, our zero-shot pipeline produces semantically labeled 3D point clouds without manual annotation, per-species fine-tuning, or domain-specific preprocessing. Multi-view NeRF fusion acts as effective implicit consensus mechanism that lifts imperfect per-frame masks into accurate 3D labels. On a controlled Begonia maculata testbed the SAM3 pipeline achieves 92.6% mIoU, reaching 95.9% of the oracle upper bound established with perfect ground-truth masks. The pipeline was further evaluated on a new dataset spanning ten diverse plant point clouds reaching an average 0.856 mIoU, with leaf and pot IoU above 0.91 and 0.90 for every species, respectively. These results demonstrate that annotation-free 3D plant organ segmentation is now feasible and approaching the range of supervised methods.
Geant4 is a toolkit for simulating particle interactions in matter that is widely used in particle physics, nuclear physics, medical imaging, and astrophysics. At present, Geant4 visualizations are generally viewed on a conventional, two-dimensional computer display. To enhance immersion, we have developed G4VR, a virtual reality (VR) application that extends Geant4 with VR capabilities, providing a more intuitive and comprehensive way to understand these interactions. The user running G4VR in a headset can get detailed information about individual tracks, interactions, and energy depositions by selecting them with the hand controllers. G4VR also allows the user to examine the chronology of interactions, and includes animations of entire events as immersive movies. G4VR includes built-in event displays of several types that allow a first-time user to easily explore its various functionalities. To interface with Geant4, we have developed a new visualization driver called G4XR that allows direct integration with the Geant4 source code. This enables any Geant4 experiment to be loaded and experienced in VR seamlessly.
Reconstructing a three-dimensional left-ventricular (LV) endocardial surface from cardiac magnetic resonance (CMR) data is challenging when supervision is available on only a small number of axial slices. Through-plane geometry is weakly constrained, and automatically generated two-dimensional masks can propagate segmentation errors into the recovered shape. We present MR-RS-SDFR, a per-case implicit signed distance field (SDF) framework that reconstructs a continuous LV surface from a CMR volume and sparse axial weak masks. The method first builds a cross-slice SDF initialization from axial and longitudinal geometric cues and then refines the field using two complementary signals: MRI edge-field normal alignment, which provides an image-derived boundary cue independent of the weak masks, and differentiable reslice Dice and contour consistency, which preserve agreement with the observed planes. We evaluate three weak-mask generators -- LOO TransUNet, LOO nnU-Net, and an off-the-shelf Medical SAM3 model used without MM-WHS-specific training or fine-tuning -- and five sparsity levels from 4 to 64 axial planes. In the sparse-16 setting, final MR-RS-SDFR reconstruction reaches 0.928 Dice and 3.80mm HD95 with Medical SAM3 masks. The upstream generators do not exhibit a single common ranking across 2D and dense 3D segmentation, and nnU-Net- and Medical-SAM3-driven sparse reconstruction achieve the same mean final Dice despite different upstream error profiles. Across all three sparse-16 mask sources, MR-RS-SDFR is numerically better than protocol-matched full GHD+DVS in both Dice and HD95. Final Dice improves markedly from sparse-4 to sparse-16 and then saturates at the reported precision through sparse-64. These results support MRI-guided per-case SDF refinement as a reconstruction strategy that remains effective across weak-mask generators and supervision densities.
Federated Learning (FL) with Differential Privacy (DP) is increasingly adopted to preserve data confidentiality in distributed machine learning. However, DP noise distorts learned representations and degrades explanation fidelity, limiting differentially private FL where trustworthy explanations are required, such as assistive clinical diagnosis. Prior work adapted DP noise with static feature-importance signals, restricting explainability to post hoc analysis and precluding noise calibration to explanation quality during training. We propose XCal-FL, a closed-loop, explainability-driven local training algorithm for image classification in cross-silo FL that dynamically calibrates DP noise from three complementary signals: (1) prediction logit variations, measuring causal influence on model confidence, (2) counterfactual margins, capturing decision-boundary sensitivity, and (3) saliency concentration, quantifying spatial coherence of model attention, while enforcing formal DP guarantees via adaptive privacy accounting. Experiments on three medical imaging datasets across varying FL configurations show that XCal-FL yields more accurate and interpretable global models, improving predictive performance by over 10\% and explanation fidelity by up to 5$\times$ over static-noise FL, and outperforming state-of-the-art adaptive DP methods in fidelity. XCal-FL also achieves higher privacy-budget efficiency, turning each unit of cumulative privacy loss into larger gains in both accuracy and explanation fidelity. Our analysis further reveals that, unlike predictive performance, which scales roughly linearly with privacy loss, explanation fidelity exhibits non-linear dynamics. These findings suggest explainability is a distinct dimension of the privacy trade-off that cannot be inferred from utility alone, with implications for training and privacy-budget allocation in decision-critical applications.
The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in whole-body PET/CT. Although recent advances in 3D medical vision-language models have demonstrated remarkable progress, current efforts are limited to regional CT imaging, leaving a critical void in comprehensive whole-body PET/CT analysis. In this work, we introduce MetaStructAtlas, a large-scale dataset for grounded whole-body PET/CT interpretation that synthesizes multimodal imaging with integrated anatomical, metabolic, and semantic annotations. MetaStructAtlas provides 490 co-registered 3D PET and CT volumes with 50,470 organ-level segmentation masks and grounded radiology reports. To facilitate interactive reasoning, we further developed MetaStructVQA, a standardized 3D grounded visual question-answering benchmark containing 100,565 QA pairs. This framework explicitly links diagnostic queries to visual evidence across modalities, encompassing anatomical, morphological, and metabolic characteristics. Finally, we evaluate state-of-the-art 3D medical VLMs on MetaStructVQA, establishing a robust foundation for multimodal representation learning and integrated whole-body reasoning in nuclear medicine.
🏫 主要单位:Division of Adult Psychiatry, Department of Psychiatry, University Hospitals of Geneva, Geneva, Switzerland; Institute of Life Sciences, Sant'Anna School of Adv | 主要作者:Beatrice A. Milano、Georgiadis, F.
Substance use disorders (SUD) are associated with widespread brain structural alterations, but the principles governing their regional distribution remain unclear. We tested whether SUD-related morphometric differences are shaped by normative connectivity and neurotransmitter receptor architecture. We analysed T1-weighted MRI data from 2,782 individuals with SUD (mean age 34.6 years [SD 11.2]; 33.3% female) and 1,951 healthy controls (mean age 33.8 years [SD 12.0]; 37.5% female) across 51 ENIGMA Addiction Working Group sites, spanning alcohol, amphetamine, cannabis, cocaine, nicotine and opioid use disorders. Cortical thickness (68 regions) and subcortical volume (14 structures) were quantified using harmonized ENIGMA protocols. Case-control maps were related to normative functional and structural connectomes through hub-vulnerability and epicenter analyses, and to 20 PET-derived receptor-density maps using partial least squares analysis. Across substances, SUD was associated with lower cortical thickness and subcortical volume involving frontal, parietal, temporal and limbic regions. Alterations overlapped with cortical hubs and converged on cortical and subcortical epicenters, with findings robust across six substances, biological sex and analytical approaches. SUD epicenters showed strong spatial overlap with those of schizophrenia and bipolar disorder, particularly in frontotemporal, parietal and striatal regions, distinct from mood, neurodevelopmental and obsessive-compulsive conditions. Cortical alteration patterns aligned with two neurochemical axes: one enriched in cannabinoid, opioid, metabotropic glutamate and histamine receptors, and another associated with lower dopaminergic and cholinergic marker density. These findings indicate that addiction-related brain differences are organized by normative connectome and neurotransmitter architecture, providing a systems-level framework for identifying receptor- and circuit-based targets for intervention in addiction.
Federated learning (FL) lets institutions train a shared model without exchanging data, and Low-Rank Adaptation (LoRA) makes this practical at scale by communicating only compact low-rank updates. Biomedical imaging is a compelling setting for this combination: patient data are archived behind privacy regulations, and institutions differ widely in scanners, protocols, and compute. Such heterogeneity raises the question of how federated LoRA updates should be aggregated, increasingly pressing as multimodal vision-language models become central to medical image analysis. We benchmark federated Parameter-efficient fine-tuning (PEFT) of BiomedCLIP for chest radiograph classification across four public cohorts on three continents (USA, Vietnam, Spain). Federated LoRA adaptation improves shared-class AUC on all four cohorts over the unadapted BiomedCLIP backbone (mean 0.687 to 0.802), showing that the gains come from federated adaptation rather than from the pretrained model's zero-shot ability. Relative to isolated single-cohort training, federation improves the weaker cohorts while largely preserving the strongest and approaches a centralized reference (0.812) that pools all data. The singular value decomposition (SVD)-based product-space aggregation introduced by FlexLoRA is essential to this gain (naive factor averaging drops mean AUC by 0.097), whereas a drift-correcting optimizer (FedProx) shows no benefit over FedAvg in our single-seed runs, consistent with LoRA's low-rank updates already limiting client drift. Biomedical vision-language models can thus be adapted collaboratively across heterogeneous, geographically distributed institutions without centralizing data.
Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration. Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real-world settings. Reinforcement learning (RL) has recently emerged as a promising alternative, requiring only the information if two images show the same scene or not. However, existing RL formulations such as RIPE rely on coarse binary rewards and carefully constructed negative training pairs, limiting training stability and descriptor discriminability. In this paper, we revisit RL-based keypoint learning and propose a reward that fully exploits the geometric consistency signal, deriving both reward and penalty from a single positive pair without contrasting against negatives. This richer signal provides sufficient supervisory contrast to learn discriminative detectors and descriptors from positive image pairs alone, enabling representation learning under extremely limited supervision. Furthermore, we show that the same RL objective can be extended to the matching stage by adapting LightGlue, raising AUC@5 on MegaDepth1500 from 56.58 to 59.65 and enabling weakly-supervised training of the full sparse matching pipeline from image pairs with partial visual overlap. We validate our approach on established benchmarks, demonstrating competitive results compared to fully-supervised methods. We further show that the method can be even trained on low texture medical video sequences, where camera poses are usually unavailable and standard SfM pipelines often fail. Code and data are available at https://github.com/fraunhoferhhi/RIPEpp .
Complex cognition depends on directed communication between brain regions. Yet, resting functional MRI resolves only undirected associations, and inferring the direction of signalling has proven difficult across the whole brain. Because synaptic transmission is metabolically asymmetric, with the postsynaptic neuron bearing most of the energetic cost, regional energy demand carries a physical signature of the direction of communication. Here, we show that this signature organizes the human cortex into a reproducible directed hierarchy, in which sensory and attentional systems preferentially drive higher-order networks, while the default mode and control networks act as principal receivers. Critically, the inferred directionality is independently predicted by two cellular markers of neuronal energy demand, regional mitochondrial density and laminar cytoarchitecture, linking macroscale information flow to its cellular substrate. The directed architecture recapitulates the known feedforward and feedback organization of visual pathways and, using an average CMRGlc template, extends to functional MRI acquired without PET. By grounding communication direction in energy metabolism, Metabolic Connectivity Mapping bridges correlative and effective connectivity and offers a scalable, biologically interpretable framework for mapping directed communication across the human brain.
Dense voxel-level annotation remains a major bottleneck in 3D medical image segmentation. Single-slice propagation methods such as Sli2Vol reduce this burden by propagating one annotated seed slice through a volume using label-free registration. However, axial-only propagation accumulates errors with distance from the seed, especially in surface-distance metrics, because it ignores coronal and sagittal evidence and therefore underuses the 3D information available in CT/MRI volumes. To better leverage volumetric geometry, we study how key training and inference choices affect slice-propagation models, including single-axis versus multi-axis label-free registration, single-seed versus multi-seed propagation, and orthogonal seed configurations. Instead of propagating from a single axial seed, we use three orthogonal seeds---one axial, one coronal, and one sagittal---and fuse their propagated labels with a simple label-free rule. Our results show that the training paradigm has limited impact: an axially trained network applied to off-axis seeds captures nearly all the improvement, while explicit three-axis training adds little. Instead, performance is driven by inference-time seed geometry, especially orthogonality rather than the number of annotated slices, as a budget-matched three-axial control provides no benefit and can even degrade performance. On a multi-organ CT cohort, orthogonal seeding with the axial Sli2Vol backbone improves Dice by 21.9%, Normalized Surface Dice by 25.5%, and reduces Average Hausdorff Distance by 53.5% over the single-axis baseline.
In practical settings, medical image segmentation models are often developed with limited annotated data rather than fully labeled datasets. Training frequently begins in ultra-low labeled regimes where only a small number of volumes are annotated. In such scenarios, practitioners must simultaneously decide which cases to annotate and how to best use the remaining unlabeled data. Although active learning (AL) and semi-supervised learning (SSL) both target annotation scarcity, they are typically designed and optimized independently, resulting in objective mismatch and unstable training during early-stage "cold start" conditions. We propose RegAL, a unified active semi-supervised framework governed by a shared topology-aware Pareto optimization that couples sample acquisition with unlabeled data utilization. RegAL evaluates images along three complementary axes, voxel-wise uncertainty, feature diversity, and a novel topological consistency metric, to select anatomically informative edge cases for annotation. On the other hand, the same criteria are used to identify geometrically stable atlas candidates for diffeomorphic registration-guided augmentation to train a self-supervised Mean Teacher segmentation network. Across BraTS 2021, dHCP, and ProstateX, RegAL remains stable with few labeled volumes and consistently outperforms state-of-the-art AL, SSL, and active semi-supervised baselines across Dice and boundary-distance (ASD, HD95) metrics under extreme annotation scarcity.
Objectives: To determine how clinically anchored MRI representations behave from supportive internal testing through locked institutional transfer and limited target-site adaptation for prediction of substantial lymphovascular space invasion (LVSI) in endometrial cancer. Materials and Methods: This retrospective two-centre study included 206 women undergoing preoperative 3-T MRI between March 2016 and March 2023. Hospital A provided training (n = 117) and supportive internal testing (n = 13); Hospital B provided strict external testing (n = 76; 12 positive). Age and CA125 formed the clinical baseline and were included in every imaging model. In 200 repeated stratified two-fold Hospital B splits, each patient was evaluated outside the adaptation half. Models remained locked; adaptation changed only the threshold or applied no-empirical-Bayes (no-EB) ComBat before threshold selection. Results: DenseNet121, U-NEXtractor, and fusion U-NEXtractor achieved internal AUC 1.000 but strict external AUCs of 0.685, 0.671, and 0.669, respectively. At development-locked thresholds, sensitivities were 0.167, 0.583, and 0.417. Threshold adaptation increased median sensitivity to 0.833 for all three, with specificities of 0.625, 0.500, and 0.500. No-EB ComBat changed median AUC by -0.007, -0.007, and -0.014, respectively, versus strict transfer. U-NEXtractor prediction ranking was most stable (median Spearman {rho}, 0.999), whereas fusion did not improve transportability. Empirical-Bayes ComBat lacked complete repeated-split coverage. Conclusion: Strong internal performance did not establish transportability. Limited target-site threshold adaptation recovered operating sensitivity without improving discrimination; feature harmonization produced no consistent AUC gain and affected representations differently.
Training deep learning-based medical image segmentation models is challenging with limited curated datasets. For AGITG TOPGEAR, a gastric cancer trial, the Clinical Target Volume (CTV) is complex and defined by multiple anatomical landmarks, making upfront training data preparation difficult for an automated contour QA segmentation model. We investigate anatomical priors, derived from surrounding organ segmentations, to provide spatial context and improve TOPGEAR CTV segmentation accuracy. We also evaluate active learning, iteratively expanding the training dataset by selecting cases expected to improve performance. One hundred TOPGEAR CT scans were retrospectively analyzed. An initial set of 10 expert-contoured cases was used to train an nnU-Net model. TotalSegmentator generated a voxel-wise anatomical prior map from surrounding structures as an additional input channel. Active learning was simulated over four iterations, selecting cases by model uncertainty and segmentation performance. All models used five-fold cross-validation for an ensemble uncertainty measure. Evaluation used a hold-out testing set of 50 cases. The anatomical prior improved CTV segmentation accuracy, increasing mean Dice Similarity Coefficient (DSC) from 0.84 to 0.86. Active learning similarly improved performance to 0.86, with greatest benefit in the final round. Combining the anatomical prior with active learning achieved the highest accuracy, with a DSC of 0.87. Model uncertainty correlated with DSC, supporting its use in identifying suboptimal predictions and guiding active learning. Anatomical priors and active learning each improved CTV segmentation accuracy and generalizability, with their combination achieving the best performance, supporting integration into segmentation model development for automated contour QA in radiotherapy clinical trials.
Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematically investigate the role of text in multimodal medical image segmentation. We first analyze several commonly used fusion strategies and find that segmentation performance is largely insensitive to the choice of fusion module. To further understand modality interactions, we propose an Evidence Decoupling Decoder (EDD) based on evidential deep learning and deep supervision. EDD serves as an internal representation analysis tool that decomposes image evidence and text-modulated evidence throughout the decoding process while maintaining competitive segmentation performance. Experimental results show that the sensitivity to text perturbation varies substantially across datasets. On BUSI and BTMRI, removing text causes catastrophic performance drops, indicating strong model reliance on textual input. On ISIC and Kvasir-SEG, text exerts relatively marginal influence. We further find that text affects predictions mainly through global semantic modulation rather than independent spatial localization, and that the specific semantic components driving text sensitivity differ across datasets. These findings provide a deeper understanding of modality interaction in multimodal medical image segmentation and offer practical insights for future model design.
Uncertainty Quantification (UQ) plays a vital role in enhancing the reliability of deep learning model predictions, especially in scenarios with high-dimensional output spaces. This paper addresses the dual nature of uncertainty -- aleatoric and epistemic -- focusing on their joint integration in high-dimensional regression tasks. For example, in applications like medical image segmentation or restoration, aleatoric uncertainty captures inherent data noise, while epistemic uncertainty quantifies the model's confidence in unfamiliar conditions. Modeling both jointly enables more reliable predictions by reflecting both unavoidable variability and knowledge gaps, whereas modeling only one limits transparency and robustness. We propose a novel approach that approximates the resulting joint uncertainty using a low-rank plus diagonal covariance structure, capturing essential output correlations while avoiding the computational burdens of full covariance matrices. Unlike prior work, our method explicitly combines aleatoric and epistemic uncertainties into a unified second-order distribution that supports robust downstream analyses like sampling and log-likelihood evaluation. We further introduce stabilization strategies for efficient training and inference, achieving superior UQ in the tasks of image inpainting, colorization, optical flow, and depth estimation.
Accurate brain tumor segmentation from magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, surgical guidance, and disease monitoring. However, developing automated segmentation models that generalize across diverse tumor characteristics, imaging protocols, acquisition sites, and patient populations remains challenging. Variations in tumor morphology and imaging distributions can substantially degrade performance outside the training domain. Consequently, improving the robustness and generalization of deep learning-based segmentation models has become a key objective in medical image analysis. To improve segmentation robustness, we propose Multi-Stage Dynamic Prompt nnU-Net, a prompt-conditioned extension of nnU-Net. Three independent dynamic prompt modules are inserted into the deepest encoder stages. Each module contains a learnable bank of ten 256-dimensional prompt vectors and uses globally pooled encoder features to generate image-specific prompt representations. These representations are projected into feature-wise scaling $(γ)$ and shifting $(β)$ parameters that modulate encoder feature maps through Feature-wise Linear Modulation (FiLM), enabling adaptive feature conditioning at multiple semantic levels. Evaluation on the BraTS GOAT validation dataset demonstrated that the proposed Multi-Stage Dynamic Prompt nnU-Net outperformed the baseline nnU Net across the majority of evaluated metrics and tumor subregions. The proposed model achieved average lesion-wise Dice scores of 76.16% (ET), 80.04% (TC), and 86.42% (WT), compared with 74.38%, 78.14% and 84.01% for the baseline model. The results demonstrate that multi-stage dynamic prompt conditioning improves segmentation accuracy and boundary delineation for brain tumor segmentation.
Objectives: Study design indexing of biomedical publications is crucial for evidence retrieval and synthesis. We sought to evaluate the accuracy and suitability of a transformer-based model (TM) for indexing clinical study designs, in comparison to National Library of Medicine (NLM) indexing. However, this is challenging for at least three reasons: First, to date, all automated systems have been trained and evaluated on manual NLM indexing assignments, itself subject to errors. Second, TM's probabilistic predictive scores take into account uncertainty, and can be converted to TRUE/FALSE assignments in different ways depending on the needs of users, while NLM labels are categorical. Third, our goal (to tag articles only that exhibit a given design) differs from NLM which tags articles that both discuss as well as exhibit that design. Materials and Methods: Therefore, we carried out a limited evaluation of the TM model that focuses only on the articles that received the most confident predictions, that is, the highest scores that are almost certainly TRUE and the lowest scores that are almost certainly FALSE, but which disagreed with NLM assignments. This was performed both for articles published in 2016 (when NLM decisions were manual) and in 2025 (when NLM decisions were automated). To establish ground truth, dual annotators indexed the articles independently, following written definitions, for four prominent study designs--cohort, case-control, cross-sectional, and case report. Results: For three designs (case-control, case report, cross-sectional), the articles having the top 100 predictive TM scores (when NLM failed to assign that design) were judged to exhibit that design in the great majority (86-100%) of cases. Conversely, the articles having the lowest 100 predictive TM scores (when NLM did assign the study design) exhibited the design only in relatively few (0-21%) of cases. The most confident predictions of the TM model were highly accurate and not redundant with automated NLM indexing; the exception was cohort studies articles, in which both TM and NLM labels showed high error rates of both omission and commission. Discussion and Conclusion: TM may have value for identifying articles exhibiting study designs, which is especially important for clinical decision-making as well as systematic reviews and other evidence syntheses. NLM indexing of cohort studies cannot be regarded as a reliable gold standard for training or evaluation of automated systems, warranting efforts to create a new manually annotated corpus.
Segmentation of curvilinear anatomical structures in 3D medical images remains challenging due to complex topology, severe class imbalance, weak contrast, and large variations in structure morphology. While deep learning approaches for 3D curvilinear segmentation have been proposed, they are often tailored to specific anatomies or modalities, limiting generalization across clinical settings and leaving room for improvement. Recent generative models have shown the benefits of iterative prediction for structured segmentation tasks, yet diffusion-based methods suffer from computationally expensive sampling, hindering their use on high-resolution 3D volumes. We present 3D-CurvSegFlow, a flow matching-based model for 3D curvilinear structure segmentation. The model learns a continuous transformation from a simple source distribution to the target vascular representation, enabling progressive refinement of complex curvilinear geometries with efficient inference. We evaluate our method on Three public challenging datasets covering distinct anatomies and modalities: portal vein, cerebral vessel, and coronary arteries. Using a common architecture and training strategy across all tasks, our method outperforms general-purpose and vessel-specific approaches, with strong preservation of thin branches and vascular continuity. This work not only advances the state-of-the-art in 3D curvilinear segmentation but also opens new avenues for efficient, generalizable, and clinically applicable methods in medical image analysis.
Diffusion-weighted MRI, beyond the commonly used diffusion tensor framework, offers a unique window into tissue microstructure in vivo, yet its clinical adoption has remained limited. Major barriers include the complexity of diffusion MRI sequence design, lengthy acquisition protocols, and the challenges associated with robust estimation of high-dimensional microstructural model parameters. Here, we address these limitations by combining optimised diffusion encoding with state-of-the-art simulation-based inference, establishing a clinically feasible framework for multi-compartment diffusion modelling. We validate the approach through i) in-depth in silico experiments and ii) in vivo studies made up of both human and rodent data. The resulting microstructural metrics are robust, reproducible across healthy individuals and show significant spatial associations with brain-wide expression patterns of cell-specific genes. Requiring less than 10 minutes of acquisition time, this framework substantially lowers the barriers to advanced microstructural imaging, a prerequisite step toward its eventual evaluation for the diagnosis, stratification, and monitoring of brain disorders.
Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity. However, whether entropy-based measures reflect clinically meaningful ambiguity, i.e. case-level disagreement about whether a pathology is present at all, remains poorly understood. Contrary to most prior work, which focused on pixel-wise boundary disagreement, we systematically evaluate how well aleatoric uncertainty captures presence ambiguity. Our evaluation spans 3D lung nodule segmentation across four architectures with Monte Carlo dropout and deep ensembles, on LIDC-IDRI and an external validation cohort (LNDb). We find that entropy-based uncertainty maps align with boundary noise and minor drawing variation but carry insufficient discriminative signal for presence ambiguity. In contrast, a lightweight supervised ambiguity head trained on frozen segmentation features substantially outperforms all entropy-aggregation-based baselines across architectures, metrics, and both cohorts, and matches or exceeds methods that explicitly model ambiguity under disagreement supervision (Probabilistic U-Net, Annotator-Confusion 3D-UNet). A qualitative feature-space analysis shows that presence ambiguity is already encoded in the frozen encoder features of pixel-wise-trained networks, only to be discarded by the segmentation output and its entropy aggregation. Our findings expose a fundamental mismatch between the theoretical promise of aleatoric uncertainty and its practical behavior, and suggest that practitioners should not rely on entropy-based uncertainty as a proxy for clinical ambiguity in safety-critical applications.
Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to exploit the rich mathematical structure inherent in medical images. We investigate whether explicit mathematical inductive biases, specifically matrix spectral analysis and vector calculus operators, can enhance segmentation beyond data-driven learning alone. Methods: We propose M-Net (Math-Augmented Network), which integrates three complementary mathematical priors into U-Net: (1) continuous spectral features derived from the condition number of centered local pixel matrices, providing a differentiable measure of texture ill-conditioning; (2) physical field operators (divergence and a discrete curl-like boundary irregularity operator) computed from image gradient fields, capturing focal intensity extrema and edge non-smoothness; and (3) a Math-Attention Gate (MAG) that adaptively fuses mathematical features with CNN-extracted deep features at skip connections. Results: Experiments on three benchmarks (LiTS, KiTS, and BraTS) show that M-Net achieves Dice scores of 78.42%, 76.15%, and 83.67%, outperforming baseline U-Net by 12.37%, 3.52%, and 5.55% on liver, kidney, and brain tumor segmentation, respectively. Ablations reveal that the condition-number feature contributes a 2.14% gain over binary invertibility features, while MAG adds 1.45% over simple concatenation. Conclusion: M-Net establishes that mathematical inductive biases provide effective complementary information for medical image segmentation. The continuous condition-number feature offers superior gradient information over discrete alternatives, and MAG preserves these priors throughout the network. This work opens avenues for integrating linear algebra and vector calculus into deep architectures for medical imaging.
Multi-modal medical images and clinical reports provide complementary anatomical, functional, and semantic information for medical image segmentation. Effectively exploiting these heterogeneous sources requires fine-grained cross-modal information fusion that preserves subtle spatial details while capturing semantic dependencies across modalities. Existing fusion approaches frequently rely on cross-attention, whose computational burden increases rapidly with spatial resolution, making dense cross-modal interaction difficult on high-resolution feature maps, particularly for volumetric medical images. In this work, we propose CoMLP, a cooperatively-gated MLP module for fine-grained cross-modal information fusion in medical image segmentation. CoMLP models cross-modal dependencies through cooperative cross-gating, built upon complementary regional and dilated MLP interactions, to capture local and global cross-modal dependencies. We further develop a multi-source fusion architecture in which CoMLP performs both inter-image fusion across imaging modalities and vision-language fusion between visual features and textual reports, enabling heterogeneous information to be integrated without relying on dense cross-attention. Extensive experiments on five medical segmentation benchmarks, covering 2D/3D images, clinical reports, multiple imaging modalities, and diverse anatomical regions, demonstrate consistent improvements over state-of-the-art multi-modal and language-guided segmentation methods. Ablation studies further show that fine-grained interaction at high spatial resolutions and complementary local-global fusion are critical to the performance gains. These results demonstrate the potential of MLP-based interaction as an effective alternative for fine-grained cross-modal information fusion in medical image segmentation.
Neuromodulatory systems project broadly, yet computations engage specific neuronal populations. Whether neuromodulation can selectively influence task-engaged neurons remains unclear. During goal-directed navigation, neurons in the locus coeruleus (LC), the brain's principal norepinephrine source, responded at navigation onset. Concurrently, dopamine transients arose in micrometer-scale domains around LC axons in CA1. These transients preferentially enhanced nearby CA1 neurons with ramping dynamics, whose higher activity predicted later reward-anticipatory action. Brief LC activation evoked local dopamine, enhanced ramping dynamics, and delayed action initiation seconds later. Blocking D1-like, but not adrenergic, receptors weakened ramping dynamics and impaired reward anticipation. Modeling showed that selectivity can emerge when localized dopamine coincides with strong ongoing neuronal activity, without dedicated wiring. Thus, a widely projecting system can selectively modulate task-engaged neurons and shape behavior.
Background: Ageing is marked by widespread cortical changes, but the molecular underpinning and network-level mechanisms associated with this variability remain poorly understood. Methods: We analysed structural MRI data from 952 adults aged 18--94 using morphometric similarity networks, subtype and stage inference, and cortical transcriptomics. Intranetwork connectivity was examined within the three cortical networks and compared with cortical expression patterns of longevity-associated genes. Results: Based on distinct intra-network morphometric similarity patterns and their associations with longevity genes, two ageing-related subtypes emerged. The normative-ageing subtype (metabolic-immune) showed connectivity profiles consistent with typical age-related decline and was enriched for genes involved in metabolism, insulin signalling, and immune regulation. The compensatory subtype (stress-repair) displayed more preserved intra-network connectivity and was associated with genes linked to stress response, DNA repair, and proteostasis. Although the two subtypes overlapped in oxidative stress and neurodegeneration pathways, they showed divergent molecular signatures associated with different cortical ageing patterns. Conclusion: By integrating network-based morphometry with transcriptomics, we establish a modelling approach that distinguishes normative decline from compensatory adaptation in cortical ageing and provides biologically informed markers of brain ageing.
Extrusion-based bioprinting enables the development of tissue-like constructs; however, the impact of printing-associated shear stress on cell viability and function remains a critical consideration. To address this, we developed a comprehensive workflow combining rheological characterization, computational fluid dynamics (CFD) modelling, and experimental validation to predict and assess shear stress effects during bioprinting. The rheological properties of gelatin methacryloyl (GelMA) at 5 % (w/v, 20 {degrees}C) and 10 % (30 {degrees}C) concentrations were modelled, comparing various non-Newtonian regression models. CFD simulations were validated using micro-particle image velocimetry, showing agreement between predicted and measured velocities. The impact of bioprinting-associated shear stress on cell viability was assessed using a co-culture of human umbilical vein endothelial cells and breast epithelial cells. Immediate post-printing analysis revealed increased apoptosis in GelMA 5 % (w/v, 20 {degrees}C), although 10 % (w/v) GelMA demonstrated higher shear stress levels compared to 5 % GelMA. After 1 day of culture in crosslinked hydrogels, apoptosis increased in extrusion pressure, demonstrating the impact of low levels of acute shear stress. This workflow provides a robust methodology for predicting acute shear stress impacts during bioprinting, laying the foundation for future optimization studies.
Machine learning models are increasingly used to predict cognitive and clinical outcomes from neuroimaging data, yet challenges in fairness and generalizability remain. Large scale datasets are often racially and ethnically imbalanced, leading to systematic performance disparities, with models typically achieving higher accuracy for majority populations represented in the training data. In this study, we evaluated whether supervised domain adaptation methods including balanced weighting, two-stage TrAdaBoost, feature augmentation with SrcOnly prediction, and linear interpolation can mitigate these biases. Using the ABCD dataset, we assessed whether models trained on 80 MRI measures from White American participants could generalize more effectively to African American participants. All domain adaptation methods reduced prediction error for African American participants, particularly for MRI modalities with large baseline disparities (e.g., structural MRI), while offering limited improvements where initial gaps were smaller (e.g., functional connectivity). Among the approaches, balanced weighting performed best and remained stable and beneficial even when only 10 African American participants were used to adapt the original model trained exclusively on White American participants. These findings suggest that simple, low-cost strategies can effectively reduce cross-ethnic performance gaps and improve equity in predictive neuroimaging, offering a practical path forward for future neuroimaging predictive biomarkers.
Memory consolidation during sleep depends on the precisely timed coordination of hippocampal reactivation with cortical slow oscillations and spindles. In humans this coupling has been characterised correlationally, while causal investigations have required intracranial stimulation in restricted patient cohorts or used non-invasive approaches targeting cortical regions, often obscuring the underlying sleep rhythms. Here we stimulated the human hippocampus during sleep for the first time, applying temporal interference stimulation to the left hippocampal head either concurrently with or prior to induced memory reactivation. Stimulation concurrent with reactivation reduced associative memory forgetting relative to mistimed stimulation and enhanced fast spindle amplitude, preserving the association between slow oscillation-spindle coupling and memory. This behavioural benefit was mediated by a distributed, left-lateralised set of spindle and coupling features. Non-invasively engaging deep hippocampal circuits during sleep offers both a means to probe human memory and a scalable basis for intervention where consolidation fails.
Multimodal image matching is a fundamental task for multi-source information fusion. However, geometric distortions and nonlinear radiometric differences (NRD) severely limit performance, especially under radiometric, rotation, and scale variations. To address this issue, we propose a radiation, rotation, and scale invariant (RRSI) feature descriptor. First, a dual-head regional sampling (DHRS) module simultaneously performs Cartesian and Log-Polar sampling on keypoint neighborhoods, retaining spatial structural properties while enhancing robustness to rotation and scale variations. We then jointly encode geometric and radiometric relations between multimodal images in a unified deep feature space, enabling feature encoding, interaction, and fusion across intra-modal, dual-head sampled, and inter-modal regions. Furthermore, we introduce a bidirectional cross-modal generative reconstruction constraint during training. By decoding implicit features into structural patches of the counterpart modality, this mechanism anchors modality-invariant geometric topologies without additional inference overhead. Experiments on optical-infrared and optical-SAR datasets demonstrate highly competitive matching performance and strong robustness to rotation and scale variations. RRSI supports the full rotation range from 0 to 360 degrees and scale factors up to four. Its generalization ability is further validated on multimodal images from computer vision, remote sensing, and medical imaging. The implementation will be made publicly available at https://github.com/yeyuanxin110/RRSI .
Four-dimensional (4D) flow magnetic resonance imaging (MRI) is a powerful non-invasive technique for visualizing and quantifying complex blood flow patterns in vivo. Despite its clinical promise, broader adoption is limited by low spatial resolution and sensitivity to noise, which restrict accurate assessment of critical hemodynamic biomarkers such as wall shear stress, pressure gradients, and turbulent kinetic energy. To overcome these challenges, we propose a deep learning-based super-resolution framework that integrates multi-scale feature extraction and attention mechanisms to enhance the quality of 4D flow MRI data. The model was trained on a dataset of 120 patients with 240 stenosed carotid arteries. High-resolution ground truth data were generated using patient-specific computational fluid dynamics (CFD) simulations based on segmented vascular geometries and physiologically realistic boundary conditions, and the resulting velocity fields served as targets for supervised learning. The proposed architecture uses convolutional block attention modules (CBAM) to guide the network toward clinically relevant spatial features and to suppress noise in low-resolution inputs. Quantitative results show that the attention-guided model substantially reduces the root mean square error (RMSE) compared with a baseline model without attention, and qualitative velocity contour analysis confirms improved reconstruction of intricate flow patterns. These findings highlight the capacity of the model to restore high-fidelity flow fields under noisy conditions and support the use of deep learning to extend the clinical utility of 4D flow MRI for non-invasive hemodynamic assessment.
High-resolution sequencing-based spatial transcriptomics, including Stereo-seq and Visium HD, aggregates dense capture units into cell-resolved expression matrices. During tissue processing and permeabilization, RNA released from source cells can spread to neighbouring capture locations, reducing cell-type specificity and biasing downstream analyses. Here we developed SPARKLE (Spatial Ambient RNA Kernel-based Leakage Estimator), a cell-level correction method that uses capture locations outside cell-segmentation masks as within-sample spatial evidence of leakage. SPARKLE fits sparse spatial kernels to out-of-mask observations to estimate a sample-level propagation scale and gene-specific leakage coefficients. It corrects only genes supported by out-of-mask goodness of fit and uses expression-dependent conservative shrinkage to protect highly expressing source cells. In ten simulated scenarios, SPARKLE achieved the highest cell-wise concordance in eight and the lowest RMSE in nine. In axolotl brain, mouse brain and human ovarian cancer, SPARKLE removed ectopic marker signal from neighbouring cells while retaining source-cell expression, improved agreement with independent single-cell and single-nucleus references, and recovered an inferred fibroblast-to-tumor COL1A2-SDC4 communication route that was obscured by ectopic COL1A2 expression. Conclusions remained stable across plausible spatial scales and background-bin sizes. Runtime scaled near-linearly with tissue-window area and was further accelerated on GPU. SPARKLE is therefore a reference-free, fast and scalable method for correcting local RNA leakage from evidence contained within each sample, improving the reliability of cell-type localization, tissue-compartment identification and cell-cell communication inference.
Vision Transformers (ViTs) have shown immense potential in medical image analysis. However, standard pre-training via global image classification suffers from spatial collapse, where models rely heavily on background shortcuts rather than localising critical foreground lesions. To overcome this limitation and align visual evidence with precise medical semantics, we systematically investigate alternative pre-training paradigms.Specifically, we evaluate three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs. Notably, our empirical analysis reveals that while all three forms of structured supervision successfully alleviate the global pooling bottleneck and steer visual attention towards foreground regions, image-text alignment achieves the most superior performance. By embedding high-dimensional diagnostic logic, the cross-modal approach not only anchors attention on precise visual evidence but also enables profound abstract reasoning. Extensive experiments demonstrate that this semantically enriched pre-training fundamentally enhances the model's feature representation. Consequently, when fine-tuned for downstream clinical classification tasks, our models achieve superior accuracy and yield highly interpretable attention maps focused on true pathological features, vastly outperforming vanilla classification baselines.
Venous thrombosis commonly develops in the vicinity of venous valve pockets, where disturbed haemodynamics, xreduced washout, and endothelial dysfunction promote thrombus initiation. While previous studies have focused on conventional flow descriptors such as velocity, shear stress, recirculation, and residence time, the influence of valve mechanics on flow organisation remains poorly understood. Here, we combine a biomimetic vein-on-a-chip platform with computational fluid dynamics, fluid-structure interaction simulations, Ghost Particle Velocimetry (GPV), whole-blood flow measurements, and particle transport experiments to investigate the interplay between valve biomechanics and venous haemodynamics. Movable venous valve leaflets fabricated by in situ photopolymerisation of poly(ethylene glycol) diacrylate (PEGDA) enabled independent control of leaflet stiffness under physiologically relevant steady and pulsatile flow conditions. Previous experiments demonstrated that leaflet flexibility governs thrombus localisation, with symmetric leaflet stiffness promoting clot formation at the valve tips, while asymmetric stiffness shifts thrombus formation towards the valve sinus. In addition, we identify a previously undescribed Reynolds-number-dependent symmetry-breaking transition in post-valve flow. Above a critical flow condition, an initially symmetric jet spontaneously develops into a stable asymmetric flow pattern. This behaviour was consistently reproduced experimentally using GPV and confirmed by Fluid-structure simulations. Valve compliance delayed the onset of the transition by increasing the effective leaflet opening, whereas valves with a small geometric offset promoted earlier asymmetry through higher local flow velocities within the valve gap. The resulting asymmetric flow generated persistent lateral bias in the transport of red blood cell-sized particles, suggesting enhanced platelet accumulation and prolonged residence within one valve sinus. These findings demonstrate that venous valve mechanics regulate not only local flow fields but also particle transport relevant to thrombosis. The discovery of a stable symmetry-breaking flow state provides a new haemodynamic mechanism linking valve stiffness, asymmetric particle transport, and the preferential localisation of thrombus formation, offering new insight into the mechanobiology of deep vein thrombosis.
While deep generative models offer new opportunities for medical image synthesis and data sharing, their ability to memorize and reproduce training samples raises serious concerns about patient confidentiality. Detecting such memorization at scale remains challenging: traditional pixel-based metrics are sensitive to generation artifacts, whereas generic embedding-based metrics often lack the anatomical sensitivity required for medical data. To address this challenge, we introduce DeepSSIM++, a self-supervised similarity metric for scalable memorization auditing in medical generative models. By leveraging multi-scale feature aggregation and anatomy-preserving augmentations, DeepSSIM++ learns an embedding space where cosine similarity approximates the Structural Similarity Index (SSIM), eliminating the need for exact pixel-level registration. Compared with state-of-the-art baselines, DeepSSIM++ achieves an average Macro F1 improvement of 33 percentage points under ideal alignment and 46 percentage points under realistic spatial and intensity perturbations. Furthermore, it accelerates large-scale similarity computation by several orders of magnitude compared with analytical SSIM. By combining anatomical sensitivity and computational efficiency, DeepSSIM++ provides an open-source tool for scalable memorization auditing in medical generative AI. Code and data are publicly available at: https://github.com/brAIn-science/DeepSSIM.
Deep learning models have achieved impressive performance in medical image diagnosis, yet their deployment in clinical settings remains constrained by limited explainability. Counterfactual images provide one means of auditing model behavior by showing how an image would need to change for a classifier to produce a different prediction. Existing approaches typically generate such explanations using auxiliary models, including generative adversarial networks and diffusion models. While often capable of producing visually realistic images, these methods explain one black-box model using another, making it difficult to separate the classifier's decision-making process from the inductive biases of the generator. We propose a novel counterfactual-generation framework that requires no generative model. Instead, counterfactuals are constructed directly from causal evidence extracted from the classifier. The resulting approach is deterministic, requires no additional model training, and enables controllable edits within user-specified regions of interest. Experiments on real-world medical imaging datasets demonstrate that the proposed method successfully changes classifier predictions while remaining closer to the original image than generative baselines, providing a more direct and transparent view of the classifier's decision boundary.
Acquiring high quality annotated medical image data is critical for training deep learning models; however, annotation is expensive, time consuming, and requires domain expertise. Conditional diffusion models, such as ControlNet, offer an alternative by generating images conditioned on semantic masks and text. However, existing approaches fail to capture fine grained properties (e.g., intensity and texture), as well as semantic consistency expected by domain experts, limiting their effectiveness for downstream tasks. Recent attempts to address these issues using reinforcement learning fine-tuning remain limited due to the reliance on a single scalar reward, which conflates diverse failure modes and provides weak corrective signals. We propose PRISM, a Compositional Reward Model (CRM) framework for conditional medical image generation. Instead of assigning a single reward, we decompose image quality into verifier grounded stages, each evaluating a distinct aspect of correctness from fine to coarse properties, including low level attributes (intensity and texture), structural alignment with conditioning inputs, and high level semantic fidelity. These stage wise rewards are composed through a Hierarchical Constrained Propagation (HCP) mechanism that enforces a fine to coarse notion of correctness, ensuring that lower level deficiencies are resolved before higher level rewards are accrued, preventing easier objectives from masking critical failures. We evaluate PRISM across three datasets spanning diverse medical imaging tasks: PanNuke (multi-class cell segmentation), CeDeM (villi/crypt detection and measurement), and ISIC (skin lesion classification). Training downstream models with data generated by PRISM yields improvements over closest baselines, including a 2.3% increase in mDice on PanNuke, a 8.5% reduction in Mean Relative Error (MRE) on CeDeM, and increases ISIC F1 by 5.9%.
Importance. Current clinical artificial intelligence is bottlenecked by highly specialized, fragmented models that reduce complex, multimodal data into isolated scalar predictions. By reframing multimodal perception as a pure vision problem, optical compression into a universal visual operating system can resolve internal data interaction bottlenecks and shift the paradigm toward continuous, generative clinical forecasting. Objective. To develop and internally validate Clinical Visual Memory, a generalist visual foundation model utilizing high-fidelity optical compression to fuse five heterogeneous intensive care unit (ICU) data modalities into a unified visual-token space, acting as a generative navigation system for critical-care trajectories. Design, Setting, and Participants. Retrospective modeling study using MIMIC-IV adult ICU encounters. After excluding stays shorter than 24 hours, 54,551 patients, 68,546 hospital admissions, and 74,829 ICU stays comprised the eligible cohort. Strict patient-level partitioning prevented cross-split leakage. Methods. Structured electronic health record (EHR) data, vital signs, 10-second electrocardiogram (ECG) waveforms, chest radiographs, and clinical notes were rendered as 2D images and encoded by a single frozen DINOv2 vision transformer. Through modality-aware latent-query cross-attention, these highly heterogeneous sources were optically compressed into a shared 1024-dimensional Clinical Visual Memory. To establish a universal output interface, this memory conditioned an conditional instruction-tuned image generator to decode eight-domain deterioration trajectories across 3- to 48-hour horizons. Visual compression fidelity was evaluated by reducing retained source-pixel area to 1%, and critical transitions were mapped using event-specific projected-axis geometry. Results. Operating as a generalist foundation, Clinical Visual Memory achieved AUROCs of 0.852 (95% CI, 0.821-0.882) for 48-hour mortality, 0.699 (0.673-0.725) for incident acute kidney injury (AKI), 0.723 (0.688-0.757) for high Sequential Organ Failure Assessment (SOFA), and 0.742 (0.726-0.757) for alive ICU discharge. Crucially, resolving the token bottleneck via an 8-fold reduction in source-pixel area preserved 98.9% of the uncompressed 48-hour mortality AUROC. Generative image-out decoding maintained clinical meaning, and trajectory geometry functioned as a proactive navigation system, yielding median warning lead times of 34.8 hours (IQR, 17.5-43.3) before death, 14.6 hours (6.5-33.2) before AKI, 15.0 hours (6.0-27.0) before high SOFA, and 20.8 hours (10.2-36.0) before recovery, with low false-alert burdens (0.040-0.109 per patient-day). Conclusions and Relevance. Clinical Visual Memory demonstrates that optical compression can successfully unify highly heterogeneous clinical data into a single, high-fidelity visual-token representation. By functioning as a continuous clinical navigation system rather than a discrete alert generator, this framework lays the architectural foundation for a generalist, generative visual operating system in critical care medicine. External and prospective validation are required before clinical deployment.
Deep learning models for CT scan analysis are often limited by the scarcity of precise pixel-level annotations, which require significant radiologist effort to produce. Training on scan-level labels alone reduces annotation requirements but introduces challenges: low supervision ratios and large input volumes make models prone to overfitting and shortcut learning. In this work, we investigate two complementary methods to address these challenges: multi-instance learning (MIL) and anatomical filtering. MIL divides CT volumes into 2D slice instances, enabling efficient 2D architectures with ImageNet pretraining rather than computationally demanding 3D models. Anatomical filtering uses Compass, our self-supervised body part regression model, to crop scans to pathology-relevant subregions without requiring segmentation masks. We evaluate two MIL frameworks - Attention-based MIL (ABMIL) and FocusMIL - on kidney tumor classification across one internal dataset (TUH) and two external datasets (KiTS23 and TCGA-KiRC). Our best models achieve F1 = 0.83 on the internal test set using only scan-level labels. We further show that anatomical filtering with the Compass model is critical for the out-of-distribution generalization of embedding-based ABMIL, while instance-based FocusMIL demonstrates greater inherent robustness to distribution shift. While evaluated on kidney tumors, we consider this a proof-of-concept for a broader weakly supervised CT classification pipeline applicable to other organs and pathologies.
Background: Glioma induces profound systemic immune alterations despite its anatomical confinement to the central nervous system. Circulating immune cells, particularly monocytes, are key mediators of tumor-host crosstalk and may retain tumor-induced transcriptional imprints. However, their potential clinical utility as blood-based biomarkers for detection and monitoring, remain largely unexplored. Methods and findings: In this study, we performed integrated single-cell RNA sequencing of blood immune cells and demonstrated that circulating CD14 monocytes are significantly expanded in glioma patients, exhibiting features of differentiation arrest and increased transcriptional plasticity. These cells harbor glioma-specific molecular signatures distinct from those observed in healthy controls and patients with other tumors. Leveraging these findings, we developed an ensemble machine learning diagnostic model based on transcriptomic profiles of circulating CD14 monocytes (training cohort, n=107), which achieved a mean area under the receiver operating characteristic curve (AUC) of 0.975 during cross-validation. In an independent cohort of 567 participants, the model maintained high diagnostic accuracy, yielding an AUC of 0.888 for distinguishing glioma from controls and other tumors. And it achieved a recurrence detection AUC of 0.975 in 51 postoperative samples. Moreover, in a follow-up study involving 30 glioma patients, lower model-derived scores of postoperation were significantly associated with prolonged progression-free survival (log-rank test, P=0.034), supporting its prognostic utility. Conclusion: We demonstrate circulating CD14 monocytes undergo glioma-specific transcriptional reprogramming, generating systemic tumor-associated signal captured via transcriptomic profiling. This blood-based diagnostic model provides non-invasive, scalable approach for glioma detection, recurrence surveillance, outcome prediction.
Brain cancer remains one of the most significant challenges in modern medicine, where the accuracy of early stage diagnosis is a decisive factor in patient survival and treatment efficacy. Although Magnetic Resonance Imaging (MRI) is the established gold standard for visualizing neurological structures, the interpretation of these high dimensional scans is often complicated by subjective variability among practitioners and the inherent noise present in complex medical images. While contemporary approaches frequently rely on high parameter deep learning architectures, such models often involve significant computational costs and require extensive data for effective training. This study introduces a hybrid framework that utilizes the Oriented FAST and Rotated BRIEF (ORB) algorithm for precise feature extraction and a Support Vector Machine (SVM) for classification [1], [2]. The proposed approach achieves a sub- stantial data reduction of approximately 99.5%, which effectively minimizes the influence of non informative background data while preserving critical diagnostic patterns essential for tumor identification. By balancing feature sparsity with a robust kernel based classifier, this methodology addresses the limitations of over parameterized systems while maintaining high diagnostic integrity. Experimental evaluations conducted on the Br35H dataset demonstrate that the framework attains a classification accuracy of 97.5%. The findings suggest that the integration of localized feature representation and optimized classification provides a reliable and resource efficient alternative for medical image analysis, offering a structured solution that maintains per- formance without the need for extensive computational overhead.
Polycystic Ovary Syndrome is a common endocrine disorder characterized by ovulatory dysfunction, hyperandrogenism, and/or polycystic ovarian morphology, with significant reproductive and metabolic consequences. Due to heterogeneous symptom profiles, Polycystic Ovary Syndrome is frequently underdiagnosed or diagnosed late. In this study, we develop machine learning models for early Polycystic Ovary Syndrome prediction using a structured clinical dataset with 42 features and 542 patient records. After data cleaning and normalization, correlation-based feature selection was applied to retain the most predictive variables. Multiple models were trained and evaluated, including Logistic Regression, Decision Tree, KNN, and Random Forest. Results demonstrate that Random Forest achieves the best overall performance (approximately 88% accuracy), suggesting that ensemble models can effectively capture non-linear feature interactions in clinical data. We also contextualize findings with international clinical guidance and recent work on explainable and clinically applicable Polycystic Ovary Syndrome prediction systems.
Background: Reliable melanoma classification requires models that capture both local dermoscopic morphology and broader contextual patterns while maintaining auditable, leakageaware internal validation. Objectives: To develop and internally validate an EfficientNetB0-Swin Transformer Tiny ensemble for classifying histopathologically verified dermoscopic images as benign melanocytic lesions or malignant melanoma. Methods: This retrospective diagnostic model-development and internal validation study screened 552,869 ISIC Archive records; filtering and dermatologist review yielded 1,199 uniquepatient and unique lesion images (578 benign and 621 malignant). Images were the predictors and histopathology was the reference. ImageNet pretrained EfficientNetB0 and Swin-T features were fused. Patient independent five fold validation used weighted sampling, mixup, label smoothing, AdamW, early stopping, and five view test time augmentation. Results: Mean accuracy was 0.89325 {+/-} 0.03179, mean receiver operating characteristic area under the curve (ROC-AUC) was 0.96348 {+/-} 0.01695, and mean support weighted F1-score was 0.89300 {+/-} 0.03220. The fold level 95% confidence intervals were 0.8538-0.9327 for accuracy and 0.9424-0.9845 for ROC-AUC. Pooled counts were 526 true negatives, 52 false positives, 76 false negatives, and 545 true positives, yielding 87.76% sensitivity and 91.00% specificity. Qualitative Grad-CAM review showed peripheral artifact activation in two false positives and lesion centered activation in two correctly classified cases; these observations were not systematically scored. Limitations: The validation folds were also used for early stopping and checkpoint selection. Device stratified analysis, systematic interpretability scoring, calibration, and independent external validation were unavailable. Conclusions: The ensemble showed high internal discrimination and is intended only as a clinician facing adjunct. The error audit workflow enables targeted retrospective review, but external validation is required before clinical use or generalizability claims.
Lesion-focused image classification presents a core analytical challenge, as discriminative signals are often sparse, spatially dispersed, and easily obscured by background noise, while conventional convolutional neural networks (CNNs) process entire images uniformly and may dilute signal relevance. This study hypothesizes that adaptive fusion of global contextual information and lesion-focused local information can improve classification performance compared with using either representation independently. We propose a three-branch, attention-guided deep learning framework built on Densely Connected Convolutional Network-121 (DenseNet-121) to improve feature attribution, interpretability, and classification reliability. The architecture consists of a global branch that learns representations from full images, followed by Gradient-weighted Class Activation Mapping (Grad-CAM) to generate attention maps that highlight prediction-relevant regions and produce masked inputs, and a local branch enhanced with a Convolutional Block Attention Module (CBAM) to extract refined spatial and channel-wise features from these focused regions. An adaptive fusion branch integrates global and local representations by learning instance-specific weights, allowing dynamic prioritization between contextual and localized information. The framework is evaluated on a synthetic Spot Pattern Dataset (SSPD) and three benchmark datasets, including skin lesion, guava leaf, and grape leaf image datasets, where the fusion branch outperformed the individual global and local branches, reaching 97.75% accuracy on the skin lesion dataset and 99.64% on the guava leaf dataset. The results highlight the value of attention-guided architectures in healthcare analytics by improving model transparency, strengthening feature relevance, and supporting more reliable data-driven decision-making in medical image analysis.
Oral potentially malignant disorders (OPMDs) are critical precursors to oral cancer, yet clinical detection remains challenging because of substantial phenotypic heterogeneity and overlap with benign conditions. Although image-based deep learning shows promise for automated screening, visual information alone may be insufficient in real-world settings, where diagnostic decisions also rely on patient-specific risk factors. We developed M2-OPMDNet, a multimodal deep learning framework that integrates co-registered white-light and autofluorescence intraoral images with structured clinical information for OPMD detection. A customized questionnaire was designed to capture clinically relevant risk factors and symptoms in a standardized, reproducible format for integration with image-derived features. Multiple image encoders, including conventional convolutional neural networks and foundation model-based architectures, were evaluated using a prospectively collected dataset reflecting real-world screening conditions. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) to quantify feature- and modality-level contributions. M2-OPMDNet achieved an AUC of 0.952, outperforming unimodal approaches and showing improved performance for visually subtle lesions. SHAP analysis demonstrated that structured clinical variables contributed substantially to risk estimation and complemented imaging features. These results demonstrate that explainable multimodal learning combining white-light and autofluorescence imaging with structured clinical data can provide accurate, transparent, and clinically grounded OPMD detection. M2-OPMDNet offers a scalable framework for real-world oral cancer screening and decision support.
Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but multi-modal benchmarks involving surgery in particular are often missing from prominent medical benchmark suites (specifically, those requiring visual recognition beyond just text question-answering). Since surgery requires coordinating disparate tasks---including multimodal data integration, human interaction, and physical effects---generally-capable AI models could be particularly attractive as collaborative tools if performance could be improved. On the one hand, the canonical approach of scaling architecture size and training data is attractive, especially since there are millions of hours of surgical video data generated per year. On the other hand, preparing surgical data for AI training requires significantly higher levels of professional expertise, and training on that data requires expensive computational resources. These trade-offs paint an uncertain picture of whether and to-what-extent modern AI could aid surgical practice. In this paper, we explore this question through a case study of surgical tool detection using state-of-the-art AI methods available in 2026. We demonstrate that even with multi-billion parameter models and extensive training, current Vision Language Models fall short in the seemingly simple task of tool detection in neurosurgery. Additionally, we show scaling experiments indicating that increasing model size and training time only leads to diminishing improvements in relevant performance metrics. Thus, our experiments suggest that current models could still face significant obstacles in surgical use cases. Moreover, some obstacles cannot simply be ``scaled away'' with additional compute and persist across diverse model architectures, raising the question of whether data and label availability are the only limiting factors. We discuss the main contributors to these constraints and advance potential solutions.
Background and Objective: External evaluation of medical-imaging AI is often collapsed into discrimination. We evaluated a computational protocol that separately tests discrimination, probability calibration, fixed operatingpoint transport, shortcut-associated signal, and limited-label recoverability for pediatric pneumonia classification across datasets from three countries. Methods: After exact-duplicate removal, 5,824 Guangzhou radiographs supported leakage-controlled source development and internal testing. A frozen three-seed DenseNet121 dual-view ensemble was evaluated zero-shot on BDCXR-3257 from Bangladesh (n = 3, 257) and an untouched harmonized VinDr-PCXR/PediCXR test cohort from Vietnam (n = 1, 077). Matched seed-42 variants tested architectural robustness. Secondary BDCXR analyses used a fixed 651-image adaptation pool and 2,606-image hold-out; 163, 326, and 651 labels represented 5%, 10%, and 20% of complete BDCXR. Results: Internal AUROC was 0.976 with 95.1% sensitivity. BDCXR and VinDr-PCXR AUROC were 0.798 and 0.742, while frozen-threshold sensitivity fell to 6.2% and 0%. Source-to-BDCXR AUROC degradation occurred for a full-image baseline (0.961 to 0.749), ungated dual-view model (0.977 to 0.766), and gated MixStyle model (0.966 to 0.789). With 163 BDCXR labels, Platt recalibration preserved AUROC while increasing held-out sensitivity to 88.3%, but specificity was 47.9% and the alert rate was 78.5%. Two hundred repeated 163-label fits confirmed sensitivity recovery but substantial specificity variability. Conclusions: Cross-dataset shifts across countries affected ranking, probability alignment, and source-defined decision behavior differently. Transport studies should evaluate these components separately and quantify the operational burden of apparent recovery.
The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetime, yet fully exploiting these data to represent and predict patient state trajectories remains a critical challenge. Current AI models often struggle to capture the complex, irregular temporal dynamics and inherent stochasticity of real-world multimodal patient data. Existing AI approaches for modeling longitudinal patient records are predominantly discriminative, limited to a few modalities, constrained by closed categorical vocabularies, treating time as a monotonic inductive bias, or they are limited in forecasting future patient states. We introduce NOAH, a time-aware, task-agnostic, generative transformer model representing and forecasting the full multimodal patient journey. NOAH features a novel bidirectional time integration and a variational latent space to capture the continuous evolution of patient states and the stochasticity of clinical trajectories. Built from over 559 million clinical events from 431,000 hospital visits of 299,000 patients across the MIMIC dataset family, NOAH natively processes medical images, time-series and numeric signals, categorical events, as well as structured and unstructured clinical records. NOAH is the first truly holistic generative model in its field, enabling autoregressive forecasting with optional time control, zero-shot classification, and counterfactual intervention simulation. It generates highly informative and predictive patient state representations that demonstrate strong performance in probing for clinical outcomes, 15 ICD chapters, and 29 comorbidities, as well as in time-to-event prediction. Seamlessly handling diverse modalities and complex temporal dynamics, NOAH provides a versatile, task-agnostic, scalable foundation for intelligent predictive systems in personalized clinical care and digital medicine.
There is growing interest in adopting CLIP-style vision--language model (VLM) pretraining for mammography. However, models that directly employ the standard CLIP architecture and training objective exhibit limited zero-shot performance in clinically important tasks such as cancer, finding-type, and BI-RADS predictions. We argue that this underwhelming performance is due to neglecting two characteristics of mammography data: (1) its high-res nature, and (2) homogeneity of radiology reports, largely driven by a predominance of negative/benign findings on examinations. We propose TopKSigLIP, a VLM designed to address these two limitations through a novel architecture and learning objectives. Instead of downscaling high-res mammography images to satisfy GPU memory constraints, TopKSigLIP introduces TopK-Patch module that learns to sample a sparse set of high-res patches likely to contain lesions, sidestepping the resolution--batch size tradeoff of VLM training. The sampled patch locations additionally serve as a built-in localization tool. To address report homogeneity, we replace the contrastive loss, which falsely repels semantically similar pairs, with a Sup-sigmoid loss. Sup-sigmoid loss extends the sigmoid loss from SigLIP with soft labels derived from structured data. TopKSigLIP outperforms existing open-source mammography and general medical VLMs on both internal and external benchmarks on density assessment, BI-RADS classification, finding subtyping, and cancer prediction under zero-shot evaluation. TopKSigLIP remains competitive under linear probing despite using a significantly smaller vision encoder and smaller training batches than baselines. The TopK-Patch module additionally achieves superior lesion localization over post-hoc Grad-CAM. Code and weights are made public:https://github.com/Youngseok0001/TopKSigLIP.
Medical AI models have made a great impact on biomedical research and real-world clinical applications, but conducting interdisciplinary medical AI research remains challenging, requiring close collaboration between clinicians and AI experts. Recent advances in large language models (LLMs) and autonomous code agents present an opportunity for low cost medical AI development, where clinicians can build AI tailored to their own research questions, even without continuous support from dedicated AI experts. However, enabling code agents to autonomously tackle complex multimodal medical AI development tasks requires clinicians to construct and supervise an AI research loop with detailed technical specifics, demanding substantial expertise in AI and computer science that they often lack. To address this challenge, we introduce the Medical AI Research Loop Agent (MARLA), an agentic framework that completely abstracts the construction and supervision of medical AI research loops from clinicians. Given a clinician-defined research intent, MARLA automatically translates high-level research goals into executable hierarchical research loops, decomposes them into verifiable sub-loops, and specifies the models, datasets, tools, and evaluation protocols required for each task. During execution, MARLA coordinates specialized code agents, monitors progress, diagnoses failures, and iteratively refines research strategies based on experimental feedback to drive the research process toward optimal outcomes. We evaluate MARLA on multimodal medical AI tasks that require closed-loop conversion from high-level clinical study objectives to trained and validated AI models. The results demonstrate MARLAs ability to autonomously conduct complex medical AI research while substantially reducing the need for AI expertise.
Developing competitive deep learning baselines for medical imaging remains a highly iterative process requiring literature review, implementation, experimentation, and expert refinement. Existing automation approaches typically optimize isolated components, such as architecture search or hyperparameter tuning, rather than the complete baseline development process. We present an agentic AI Scientist workflow that combines literature-guided reasoning, automated code generation, and hypothesis-driven experimentation to generate competitive baseline models for medical imaging challenges. The framework is evaluated on four public benchmarks spanning segmentation, classification, and detection. Across all tasks, the Experimentation Pipeline consistently improves validation performance, achieving competitive leaderboard results, including 6th place on both PUMA tracks (15 teams) and 31st place on MILK10k (125 teams). On MIDOG25, the resulting model also demonstrates strong domain generalization across scanners, tumor types, and species. Using the same workflow across all challenges without task-specific redesign, we demonstrate that skill-based, literature-guided agentic workflows can substantially reduce the engineering effort required to develop competitive medical imaging baselines.
Vision-language models (VLMs) are increasingly deployed in high-stakes settings, where a response that is reasonable in general may still be unsafe for a particular user whose medical, emotional, or situational context is unknown to the model. We study this problem of personalized safety in multimodal systems and introduce MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile. Evaluating eight frontier VLMs, we find that they almost always respond directly (86-99%) rather than seek missing context, and none exceeds 2.6/5 on personalized safety. To understand why these failures arise, we analyze multimodal interactions and identify visual dominance: visual information enters text representations early and suppresses textual risk signals during multimodal fusion. Causal interventions reveal a two-stage mechanism in which visual affect is first transferred into the text stream in early layers and then shapes the final decision through this altered text representation, making late-stage internal remediation unreliable. Motivated by this mechanism, we propose PRISM, a lightweight input monitor that uses bidirectional cross-modal modulation to predict when a query is likely to require deferral. PRISM achieves 0.978 AUC and strictly dominates the safety-utility Pareto frontier across all tested models.
Medical vision-language models (VLMs) require confidence that reflects both answer correctness and patient-specific visual evidence. Recent GRPO-based methods optimize verbalized confidence together with answer generation. However, this joint optimization may interfere with answer learning and drive confidence toward near-binary values. Verbalized confidence also provides no explicit assessment of visual support. We therefore separate capability learning from confidence estimation and propose \textbf{DualRead}. DualRead builds on the insight that reliability can be read from the actor's internal states at critical moments in the answering process. It freezes the GRPO-trained actor and combines pre-answer solvability with a post-answer assessment of the generated answer and its visual support. To further assess whether confidence reflects visual grounding, we introduce \textbf{Counterfactual Confidence Grounding AUC} (CCG-AUC). It measures whether confidence decreases when real-image substitution changes the actor from correct to incorrect. Across two VLM backbones and both in- and out-of-distribution medical VQA benchmarks, DualRead improves correctness discrimination and calibration over verbalized confidence while preserving answer accuracy. CCG-AUC reveals whether confidence responds to answer-relevant visual evidence rather than primarily to non-visual cues.
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of clinically reliable and linguistically inclusive medical AI systems remains a significant challenge, primarily due to the lack of multimodal, multilingual, and time-series-grounded benchmarks that reflect the complexity of real-world clinical scenarios. To fill this gap, we present MMTClinic, a benchmark designed to evaluate large language models (LLMs) on complex reasoning and question-answering tasks involving clinical time-series. MMTClinic combines text, medical images, and multivariate physiological signals and includes 30,000 QA pairs (15,000 multiple choice questions (MCQs) and 15,000 open-ended questions) across five languages: English, Hindi, Bengali, Marathi, and Tamil. These questions cover three important clinical tasks---mortality prediction, heart rate forecasting, and SOFA score estimation. We evaluate 13 state-of-the-art LLMs in zero-shot, few-shot, and chain-of-thought settings. Our evaluation reveals notable differences in model performance across tasks, languages, and modalities, highlighting current limitations in clinical reasoning capabilities. MMTClinic provides a valuable resource for advancing multilingual, multimodal, and time-series-aware medical AI research. The dataset will be made publicly available on successful acceptance of the work.
Purpose: To develop and optimize a 5DCT (3D + cardiac phase + respiratory phase) imaging simulation and reconstruction pipeline, and to compare two sinogram-space interpolation methods for reconstructing images at arbitrary combinations of cardiac and respiratory phase. Methods: Helical CT projections were simulated from the 4D XCAT phantom across a range of cardiac and respiratory motion states, with Poisson and electronic noise added. Ground-truth-matched volumes were generated at 5 cardiac phases and 10 respiratory amplitudes (50 total phase combinations). Because acquired projections are sparsely and unevenly distributed across this joint phase space, each target slice was reconstructed by interpolating rebinned sinogram rows to the target cardiac phase and respiratory amplitude, using either 2D scattered barycentric interpolation or 2D scattered local linear interpolation with a circular kernel for cardiac phase. Reconstructed volumes were compared to phantom ground truth using mean absolute error (MAE), and to conventional respiratory-gated 4DCT (r4DCT) reconstructed from the same simulated data. Results: Both interpolation methods eliminated the severe axial misalignment artifacts present when helical projections were reconstructed without phase-space interpolation. Local linear interpolation achieved lower MAE than barycentric interpolation across most tested conditions, with the largest improvement at low pitch. The 5DCT pipeline also produced respiratory-only volumes with fewer residual cardiac-motion artifacts than conventional r4DCT reconstructed from the same projection data, including at standard clinical pitch (0.1). Conclusions: 5DCT reconstruction using sinogram-space interpolation is feasible and can jointly resolve cardiac and respiratory motion with better accuracy than conventional 4DCT reconstruction.
Purpose: We introduce PRE-CISE, a pre-calibration workflow integrating coverage analysis, local sensitivity, and collinearity diagnostics to streamline model calibration. We demonstrate PRE-CISE's benefits using two testbeds (a Sick-Sicker Markov model and an SIR transmission model) and a COVID19 case study. Methods: PRE-CISE uses coverage analysis to verify that model outputs generated from the prior distribution span calibration targets, followed by local sensitivities to quantify parameter influence, guiding the resizing of prior bounds to improve coverage. Identifiability is assessed via collinearity analysis; large indices indicate practical nonidentifiability. Bayesian calibration was used: for the testbed models (3 and 2 parameters, respectively) to match their targets; for the COVID-19 model (11 parameters) to match daily confirmed incident cases. Results: Coverage analyses flagged initial misfits; local sensitivities identified that the Sick-to-Sicker transition probability has a greater effect on model outputs, and resizing its prior distribution bounds improved coverage. Collinearity analyses indicated that using multiple calibration targets over different time points enabled identification of all three parameters. For the SIR testbed, collinearity indices declined as target points accumulated, crossing the identifiability threshold once the data spanned the epidemic peak. In the COVID-19 model, local sensitivity analyses prioritized time-varying detection rates and contact-reduction effects, reducing the search space. Daily incident case calibration targets yielded collinearity indices below practical thresholds for all parameter combinations, whereas indices for weekly calibration targets were larger. PRE-CISE achieved greater calibration efficiency gains for denser posterior sampling and more complex models. Conclusions: PRE-CISE provides a practical, transparent pathway that helps modelers refine prior distribution bounds and calibration targets before intensive calibration, improving uncertainty reporting and strengthening the reliability of model-based health policy analyses.
Deformable image registration models implicitly encode deformation priors through their parametrization and optimization. In this work, we conduct a validation study on continuous registration methods to examine how these implicit priors affect performance across different registration tasks. Classic B-Spline transformations impose locality, smoothness, and scale through their control-point structure, whereas recent INR-based methods impose different priors through neural parameterization and optimization. We compare INR-Dense (IDIR), which directly models a dense displacement field using a SIREN-based INR; INR-BSCP (SINR), which predicts B-Spline control points with an INR; D-BSCP, which directly optimizes single-scale B-Spline control points; and MR-D-BSCP, which adds a multiresolution coarse-to-fine scheme. Experiments on inter-subject brain MR registration (OASIS) and intra-subject exhale-to-inhale lung CT registration (DIR-LAB 4DCT) reveal different behavior across deformation regimes. On OASIS, where deformations are moderate but locally complex, D-BSCP matches or slightly outperforms INR-BSCP, suggesting that the B-Spline parameterization accounts for much of INR-BSCP's effectiveness. On DIR-LAB 4DCT, where respiratory motion is larger and more coherent, single-scale B-Spline methods (D-BSCP and INR-BSCP) are less suitable, while INR-Dense and MR-D-BSCP are more effective. Across both tasks, MR-D-BSCP achieves the best performance among the tested continuous parameterizations. These findings highlight that registration accuracy depends strongly on matching the induced deformation prior to the target motion pattern, and support prior-deformation matching as a practical design principle for medical image registration. Our code will be available at https://github.com/HengjieLiu/RightPriorDIR.
Registration-based Few-shot medical image segmentation (RFMIS) aims to generate pseudo-labels for unlabeled images by warping a labeled image through registration. However, existing methods primarily perform pixel-level optimization and inference in Euclidean space, treating anatomical structures as flat and disjoint. This neglect of inherent hierarchies degrades pseudo-label quality and weakens the discrimination of ambiguous regions, limiting the segmentation performance. To overcome this challenge, we propose a Hyperbolic Hierarchy-aware Aggregative Learning framework for RFMIS, termed H2AL, that enhances both deformation plausibility and anatomical discrimination for dual-task learning. Specifically, we introduce a Hyperbolic Hierarchy-aware Infusion (H2I) module, which leverages the hierarchical modeling capability of hyperbolic space to learn precise hierarchy-aware representations via transformation-guided supervised hyperbolic contrastive learning, and injects such hierarchical priors into Euclidean space through a gated infusion block while preserving semantic richness. Furthermore, we propose an end-to-end joint optimization algorithm by gradient aggregation, where the gradients from the registration and segmentation decoders, embedding semantic and hierarchical cues, are aggregated to update the shared encoder to promote collaborative learning across tasks. Extensive experiments on two anatomical regions, with five experimental settings, demonstrate the effectiveness and efficiency of our method in both registration and segmentation. The code is publicly available at https://github.com/JiamingCai469/H2AL.
Deformable image registration (DIR) is a core problem in medical image analysis; but, unlike labeling decision problems such as classification and segmentation, registration is a problem class that involves stringent physical constraints. Although deep learning methods have made faster registration possible, the resulting models are often difficult to interpret compared to hand-crafted methods with explicit objectives and interpretable physical meaning. In this work, we show that an analytical method can still yield competitive and superior results to deep learning in a common deformable registration task. We study pTVreg as a parametric total variation based registration in that context. Observing its different implementations to perform at various degrees, we introduce here an accessible implementation of this method, together with a Bayesian optimization framework that automatically sets self-parameters for any DIR task from a set of sample examples. Experiments on Lung250M-4B show that our proposed implementation achieves state-of-the-art results in this benchmark, substantially superior to existing deep learning solutions and other pTVreg variants as baselines. The source code will be made publicly available at https://github.com/oazeybekoglu/ptvreg-python .
Head-and-neck (HN) proton therapy is highly sensitive to anatomical change over a 4-to-6-week course, as tumor shrinkage, weight loss, and setup variation can misposition the Bragg peak near critical organs such as the parotids, oral cavity, brainstem, and spinal cord, leading to target underdosing or organ-at-risk overdosing. Online adaptive proton therapy replans on the anatomy of the day, yet standard workflows rely on offline replanning that requires repeated CT acquisition and roughly a week of preparation, adding burden, cost, and delay. We investigate whether a patient's treatment-day anatomy can be predicted before image acquisition by transferring longitudinal change from a population database. We propose a digital-twin framework built on a pretrained foundation-model deformable registration network used without patient-specific training. A first registration aligns a prior patient's planning CT to the target and carries the prior's during-treatment quality assurance CT (QACT) into the target frame; a second registration estimates the prior's planning-to-QACT change, which is then applied to the target's own planning CT to synthesize predicted CTs (pdCTs) with propagated contours. Using 88 HN patients, each with a planning CT and three QACTs, we show that pdCTs better match treatment-day anatomy than the static planning CT. Compared with the planning CT alone, normalized cross-correlation improves by 22.8%, Dice for organs-at-risk by 20.2%, and CT-number error decreases by 23.4%. Gains are largest for patients with major anatomical change and negligible when anatomy is stable. This cross-patient motion transfer leverages the digital-twin concept to anticipate treatment-day anatomy, enabling personalized online adaptive proton therapy without repeated imaging.
Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision. Existing landmark-free approaches perform partial-to-complete geometric alignment, yet robust self-supervision under extreme partial visibility remains difficult. We propose Vis2Reg, a visibility-aware registration framework that explicitly constrains deformation using mask-consistent visible regions. We introduce a visibility-aware self-supervision that derives a visible-domain 3D supervision signal from intraoperative masks, enabled by differentiable point rasterization and mask-guided back-projection. This formulation improves robustness under severe occlusion while maintaining fully self-supervised learning. Vis2Reg combines a robust geometric rigid initialization module with an implicit neural deformation field for stable alignment. Vis2Reg achieves a Dice score of 92.6\% and a Chamfer Distance of 1.43 mm on real intraoperative datasets, with 111 ms per-frame inference time, demonstrating both accuracy and practical efficiency.
Strongly enforced Dirichlet boundary conditions require highly refined near-wall meshes to resolve steep velocity and thermal gradients. This introduces high computational costs, especially for practical flow simulations. Weakly enforced boundary conditions alleviate this burden by acting as a variationally consistent near-wall model. By allowing a controlled slip at the solid wall, weak enforcement recovers accurate flow quantities on coarse boundary-layer meshes across both incompressible and compressible regimes. Furthermore, weak boundary conditions serve as the fundamental enabling technology for immersogeometric analysis. Because the weak operator evaluates boundary integrals independently of the background mesh, high-fidelity flow analysis can be performed directly on complex geometries without fitting a mesh to the surface. This capability has facilitated direct geometry-to-analysis workflows for boundary-representation CAD models, raw point clouds, photogrammetric reconstructions, and segmented medical images. This review examines the unified mathematical development of the weak boundary condition framework from scalar advection-diffusion equations to the full Navier-Stokes equations, illustrating its versatility and robustness across incompressible and compressible flows, whether in traditional boundary-fitted, sliding-interface, or advanced immersogeometric applications.
To enable accurate and rapid photon control point and proton beamlet dose calculation in the DoseRAD2026 challenge, we present BEAM3R, a dose estimation framework operating in beam's-eye-view (BEV). Our core innovation combines a Mamba-3 state-space depth-sequence core with physics-based transport conditioning to model long-range depth transport without expensive 3D convolutions. BEAM3R shares a 2D CNN encoder-decoder architecture for photon and proton dose tasks, processing per-plane BEV slices. Proton beamlets are conditioned on water equivalent thickness and remaining range, encoding the parameters determining Bragg peak position. Photon models use a bidirectional Mamba-3 core to capture dose contributions from materials downstream of the calculation point, while the proton model uses a forward core with learned energy-prefix tokens and a Bragg-peak refinement module. To reduce interpolation artifacts and support high spatial resolution, we introduce axial grid alignment of BEV lattices with CT slices and an implicit super-resolution representation via sub-pixel phase packing, evaluated by a differentiable Triton-accelerated resampler that reconstructs packed cubic B-spline coefficients directly in CT space. For MRI-based tasks, synthetic CTs (sCT) are generated by a patch-based conditional GAN with a SwinUNETR backbone. On the preliminary DoseRAD2026 test set, CT-to-photon and CT-to-proton models achieved 1%/1 mm local gamma pass rates of 96.8% and 96.0%, with stratified plan-level MAEs of 0.0041 and 0.0079. Substituting sCT reduced gamma pass rates to 89.7% for photon and 75.4% proton plan level doses, with stratified plan-level MAEs of 0.0093 and 0.0336. Standardised runtimes were 23.4 s and 18.4 s for CT-to-photon and CT-to-proton prediction, increasing to 39.7 s and 42.8 s for the corresponding MRI-based pipelines.