SGP-TTA: Skeleton-Guided Progressive Test-Time
Adaptation for Thin Curvilinear Structures

1Seoul National University  2Seoul National University Hospital  3Yonsei University  4Korea University
*Equal contribution  ·  †Corresponding author
Preprint, 2026 · Under review
SGP-TTA overview

Overview of SGP-TTA. A source-trained network with frozen convolutions adapts only at its BN layers, through two complementary mechanisms. ProgBN blends frozen source and current target statistics under a sample-count schedule, without gradient updates; CSR recalls the skeleton of the multi-view consensus map and updates only the BN affine parameters.

Abstract

Accurate segmentation of thin curvilinear structures is vital for various real-world applications, from vessel analysis to road extraction. Yet their intricate geometry makes even minor pixel-wise errors enough to break the global topology, and this structural fragility turns severe domain shifts into catastrophic failures. The difficulty is most acute under cross-modality gaps, where the imaging process itself differs fundamentally between source and target. While test-time adaptation (TTA) offers a practical source-free remedy, existing methods adapt feature statistics and confidence — neither of which constrains connectivity — and thus degrade under such extreme gaps. To address this, we propose Skeleton-Guided Progressive Test-Time Adaptation (SGP-TTA). Progressive Batch Normalization (ProgBN) shifts normalization from frozen source statistics toward current target estimates under a sample-count schedule, so that the source–target balance follows the stage of adaptation rather than a fixed coefficient. Consensus Skeleton Recall (CSR) then derives a structural target from geometrically aligned multi-view predictions and updates only the BN affine parameters to preserve connected structures. Extensive experiments show that SGP-TTA consistently outperforms existing TTA methods in topological connectivity, with the largest margins under cross-modality shift.

Why cross-modality matters

Curvilinear datasets and their domain characteristics

Representative curvilinear-structure datasets and their domain characteristics. (a) Color fundus photography (CFP), fluorescein angiography (FA), OCT angiography (OCTA), and aerial road datasets share the same underlying topology but differ sharply in appearance. (b) t-SNE of batch-normalization statistics; colors denote datasets, marker shapes denote imaging domains. Modality clusters are well separated — which is exactly why source-only models collapse under modality transfer.

Method

1. Progressive Batch Normalization (ProgBN)

The affine parameters of each BN layer remain calibrated to source-normalized features. Replacing the source statistics immediately therefore destabilizes early predictions, while a fixed source-heavy mixture becomes increasingly conservative after the affine parameters have adapted. ProgBN resolves this stage-dependent trade-off by progressively increasing its reliance on the target statistics as the stream proceeds.

$$\tilde\mu_n = \alpha(n)\,\mu_t(x_n) + \bigl(1-\alpha(n)\bigr)\,\mu_s, \quad \tilde\sigma^2_n = \alpha(n)\,\sigma^2_t(x_n) + \bigl(1-\alpha(n)\bigr)\,\sigma^2_s$$

where the mixing coefficient follows a bounded square-root schedule in the target-sample count n:

$$\alpha(n) = \frac{\sqrt{n}}{\sqrt{n} + \tau}$$

The schedule is monotonically increasing and label-free. The source statistics act as a prior worth τ units of adaptation progress, giving τ a concrete pseudo-count reading rather than an opaque tuning constant; the half-transition point is n = τ2.

2. Consensus Skeleton Recall (CSR)

ProgBN recalibrates feature statistics but does not constrain the spatial connectivity of thin structures. CSR derives a self-supervised structural target from prediction agreement across K geometric views (identity, two flips, three 90° rotations). After mapping each view's prediction back to the original coordinate frame, we form a Multi-View Consensus map $\bar p_n$, binarize it, skeletonize it, and dilate it into a tube $\tilde s_n$.

CSR then encourages each aligned view to cover the consensus-supported skeleton:

$$\mathcal{L}_{\text{CSR}} = 1 - \frac{1}{K}\sum_{k=1}^{K} \frac{\sum_i \tilde s_{n,i}\, p^{(k)}_{n,i}} {\sum_i \tilde s_{n,i} + \varepsilon}$$

Combined with a pixel-wise entropy term on $\bar p_n$, a single Adam step per image updates only the BN affine parameters ($\gamma, \beta$). All non-BN weights and the source running statistics stay frozen.

Qualitative Results

Qualitative comparison across target domains

Qualitative comparison across diverse target domains. Red boxes highlight restored continuity of fine vessel structures and reduced false negatives in SGP-TTA (Ours) compared with the source model and competing TTA baselines.

Quantitative Results

Dice (%) and clDice (%) across three distribution-shift regimes. Higher is better. Best and second-best per column.

Within-modality retinal transfer

Method DRIVE → STARE DRIVE → CHASEDB OCTA3mm → ROSE1 OCTA6mm → ROSE1 ROSE1 → OCTA3mm ROSE1 → OCTA6mm
DiceclDice DiceclDice DiceclDice DiceclDice DiceclDice DiceclDice
Source51.3248.5537.0838.9547.4741.7556.4949.7259.7664.2268.9280.51
TENT70.9668.6671.6073.5148.4741.6553.7748.7963.1766.0771.7180.89
CoTTA71.6769.6671.2773.4347.5442.7553.9448.9562.2664.5271.4180.30
SAR70.9668.6971.3773.2248.5541.7353.8448.8762.1764.4471.3680.26
EATA71.0668.7871.6173.5648.5743.7353.8749.7962.4465.0571.4280.60
MedBN70.2668.1570.1371.3350.1943.8956.5451.8149.8650.8361.7671.54
VPTTA70.9268.0371.3373.4749.2342.7556.0551.0654.4355.7567.9778.44
GraTa70.9668.6971.4173.2648.5641.7453.8548.8962.3764.6671.4480.33
TopoTTA71.9870.6671.7076.1249.4443.3555.1548.5762.1865.5565.7679.23
SGP-TTA (Ours) 74.1771.36 73.7978.30 58.7252.02 62.7158.41 63.5566.30 72.5681.95

Cross-modality retinal transfer

Method DRIVE → FA19 DRIVE → OCTA3mm DRIVE → OCTA6mm DRIVE → ROSE1
DiceclDice DiceclDice DiceclDice DiceclDice
Source40.7540.0925.8424.3937.9236.6046.3544.31
TENT41.5841.8430.1433.9054.6763.0240.7442.12
CoTTA41.4241.9229.4833.0752.6460.5940.6242.03
SAR42.6142.8929.5833.5053.4261.3741.6343.25
EATA41.6041.8830.4934.0154.9162.9840.8042.18
MedBN39.9741.1827.0829.6048.6254.9041.8643.00
VPTTA44.7945.0832.5033.9155.1359.3844.3044.90
GraTa41.6241.8929.4833.1052.8860.9040.6442.03
TopoTTA43.8544.0124.1823.3335.5135.3646.6847.04
SGP-TTA (Ours) 47.5349.65 44.8346.41 59.6269.71 53.9551.16

Cross-dataset road transfer

Method DeepGlobe → MR DeepGlobe → CNDS
DiceclDice DiceclDice
Source37.2344.9485.5693.39
TENT37.1846.0562.6777.06
CoTTA37.3546.3262.9177.20
SAR37.3546.3062.9477.22
EATA37.2646.2863.1177.41
MedBN35.2742.6467.2682.41
VPTTA37.6746.8571.5284.63
GraTa37.4046.3863.0077.30
TopoTTA42.2453.5183.0391.28
SGP-TTA (Ours) 40.8054.84 85.2794.30

SGP-TTA achieves the highest clDice in all 12 transfer settings and the highest Dice in 10 of 12. Gains are largest under cross-modality shifts (Table 2), where several existing TTA methods degrade below the source model.

Analysis

Effect of the alpha schedule over the target stream

The right source–target balance shifts over time

No fixed source–target coefficient wins throughout the stream. Source-heavy normalization leads early; target-heavy normalization starts lowest and only overtakes later. The square-root schedule used by ProgBN gives the largest improvement at every point of the stream, directly supporting the stage-dependent view.

Robust across BN-based architectures

SGP-TTA is the outermost method on every axis of UNet, UNet++, MaNet, LinkNet, and FPN, averaged across all transfer settings. The method operates purely on BN layers, so the gain does not rest on a particular decoder design — and, by the same token, the claim is scoped to normalization-based networks.

Robustness across segmentation architectures
clDice vs adaptation cost

Headline clDice at reasonable cost

SGP-TTA is slower than the entropy-driven baselines because it runs multiple views and a skeletonization step per image, but still sits well below the most expensive compared method (CoTTA) while attaining the highest clDice. Entropy-driven methods cluster near the same clDice regardless of whether they update only BN affine parameters or the whole network — the limit is the objective, not the compute.

BibTeX

@article{jang2026sgptta,
  title   = {Skeleton-Guided Progressive Test-Time Adaptation for Thin Curvilinear Structures},
  author  = {Jang, Boa and Lee, JunGyu and Lee, Gwanho and Choi, Jinwook and Kim, Young-Gon},
  year    = {2026},
  note    = {Under review}
}