A meta-learned optimizer for 3D Gaussian Splatting that reaches higher quality than Adam at equal wall-clock time and keeps improving over long optimization horizons, with one model for both sparse- and dense-view settings.
Learn2Splat replaces Adam with a meta-learned network predicting per-Gaussian updates from Adam-normalized gradients and maintained latent states. The core architecture module is a kNN-based Point Transformer that captures spatial Gaussian relationships.
Stores intermediate scene states from previous meta-iterations. Exposes the optimizer to states from large early gradients through fine late-stage refinements without extending the computational graph.
Scenes are further optimized with a frozen snapshot before buffering. Rollout horizon grows from 1→50 steps over the first 10k meta-iterations, teaching the optimizer to recover from its own mistakes.
Predicts per-Gaussian scaling coefficients from Adam-normalized gradients, restoring magnitude information suppressed by transformer normalization so updates decay naturally as the loss decreases.
Trained on low-resolution DL3DV scenes only, a single Learn2Splat model generalizes zero-shot to new datasets, resolutions and view counts: in the sparse setting to high-resolution and 32-view DL3DV and to RealEstate10K, and in the dense setting to DTU, LLFF and Mip-NeRF360.
A single model is meta-trained jointly on two configurations of DL3DV, alternating between them across meta-iterations. Both use low-resolution images (256×448), with 8 context views at each inner iteration and a fixed set of 6 target views to evaluate the update; training runs for 150k meta-iterations on 4 NVIDIA A100 GPUs (about 3 days). We compare it to dedicated per-configuration checkpoints in the analysis.
A sparse-view, forward-facing setup with ReSplat feed-forward initialization. The same 8 context views, taken from a short window of the scene video, are used at every inner step, and latent states are initialized from the feed-forward features.
A dense-view, large-baseline setup with SfM point cloud initialization and random latent states. Each scene draws a pool of 64 views by farthest-point sampling, and every inner step samples 8 context views from it. Augmented by keeping 10–100% of the SfM points, with COLMAP's black points removed.
A single Learn2Splat checkpoint is evaluated on eight benchmarks: four dense-view datasets from an SfM initialization and four sparse-view configurations from a ReSplat initialization. Every optimizer gets the same runtime, Adam's wall-clock time for 4000 iterations, and runs as many of its own iterations as fit within it. All optimizers use spherical harmonics of degree 3 from the start, and adaptive density control is disabled to compare them at a fixed capacity.
| Optimizer | LPIPSβ | SSIMβ | PSNRβ | PSNRTβββ | Iters | Memβ[GB]β | |
| Dense-view, SfM init | |||||||
| DL3DV ~300 views 256Γ448 | G3R⋆ | 0.218 | 0.838 | 25.81 | 25.12 | 1608 | 1.6 |
| Adam | 0.185 | 0.879 | 27.78 | 27.78 | 4000 | 1.5 | |
| L2S | 0.157 | 0.896 | 29.04 | 29.04 | 1713 | 1.5 | |
| DTU ~50 views 1162Γ1554 | G3R⋆ | 0.350 | 0.869 | 27.19 | DIV | β | 1.0 |
| Adam | 0.309 | 0.899 | 28.96 | 28.96 | 4000 | 1.0 | |
| L2S | 0.308 | 0.900 | 29.25 | 29.25 | 1919 | 1.0 | |
| LLFF ~20β60 views 756Γ1008 | G3R⋆ | 0.355 | 0.723 | 22.77 | 22.68 | 1975 | 1.4 |
| Adam | 0.354 | 0.731 | 22.82 | 22.82 | 4000 | 1.3 | |
| L2S | 0.303 | 0.767 | 24.56 | 24.56 | 2855 | 1.3 | |
| Mip360 100+ views 520Γ780 | G3R⋆ | 0.423 | 0.658 | 23.62 | 22.68 | 1441 | 4.4 |
| Adam | 0.372 | 0.715 | 25.62 | 25.62 | 4000 | 3.4 | |
| L2S | 0.367 | 0.715 | 25.75 | 25.75 | 2013 | 3.5 | |
| Sparse-view, ReSplat init | |||||||
| DL3DV 8 views 256Γ448 | ReSplat | 0.144 | 0.854 | 26.59 | 5.99 | 1252 | 1.5 |
| G3R⋆ | 0.093 | 0.916 | 30.34 | 27.93 | 1688 | 0.9 | |
| Adam | 0.097 | 0.909 | 30.02 | 29.91 | 4000 | 0.6 | |
| L2S | 0.090 | 0.917 | 30.53 | 30.53 | 1924 | 0.9 | |
| DL3DV 8 views 512Γ960 | ReSplat | 0.325 | 0.722 | 22.08 | DIV | β | 20.2 |
| G3R⋆ | 0.192 | 0.846 | 26.99 | 23.83 | 1697 | 3.0 | |
| Adam | 0.197 | 0.837 | 26.60 | 26.32 | 4000 | 1.8 | |
| L2S | 0.185 | 0.846 | 27.16 | 27.16 | 2063 | 3.0 | |
| DL3DV 32 views 256Γ448 | ReSplat | 0.258 | 0.743 | 21.33 | 5.51 | 1797 | 12.0 |
| G3R⋆ | 0.100 | 0.911 | 29.86 | 26.75 | 2279 | 4.2 | |
| Adam | 0.103 | 0.904 | 29.78 | 29.70 | 4000 | 4.1 | |
| L2S | 0.095 | 0.913 | 30.26 | 30.26 | 2334 | 4.2 | |
| RE10K 8 views 512Γ960 | ReSplat | 0.338 | 0.745 | 19.90 | DIV | β | 20.2 |
| G3R⋆ | 0.214 | 0.855 | 26.01 | DIV | β | 3.0 | |
| Adam | 0.223 | 0.844 | 25.57 | 25.26 | 4000 | 1.8 | |
| L2S | 0.214 | 0.852 | 26.08 | 26.08 | 1818 | 3.0 | |
Optimization against wall-clock time, dense setting. PSNR averaged across scenes, plotted against elapsed optimization time. Dots and labels mark where each run reaches Adam's peak PSNR, and each panel lists the methods' best PSNR. A cross marks where a run's Gaussian parameters diverged to non-finite values. L2S reaches every quality level faster than both Adam variants and ends higher, except on DTU, where Adam-tuned ends marginally above it.
Optimization against wall-clock time, sparse setting. PSNR averaged across scenes, plotted against elapsed optimization time. Dots and labels mark where each run reaches Adam's peak PSNR, and each panel lists the methods' best PSNR. A cross marks where a run's Gaussian parameters diverged to non-finite values. L2S reaches every quality level faster than both Adam variants and ends higher, despite a more expensive update.
On 140 DL3DV scenes (in-domain) and 15 DTU, 8 LLFF and 9 Mip-NeRF360 scenes (zero-shot), Learn2Splat improves over Adam by 0.1–1.7 dB within the fixed runtime. G3R⋆ peaks within 100 steps, as it is tuned for early performance, then degrades, ending 1.8–3.2 dB below Learn2Splat and diverging on DTU.
On 140 low- and high-resolution DL3DV scenes and 72 RealEstate10K scenes, Learn2Splat stays above Adam throughout and improves by close to 0.5 dB, while Adam drops 0.1–0.3 dB below its own peak. G3R⋆ ends 2.6–3.5 dB below on DL3DV and diverges on RealEstate10K; ReSplat improves only for the first ~10 iterations.
On the nine Mip-NeRF360 scenes, in the regime of LMRS (SH degree 0, L2 inner loss, 16 views), Learn2Splat attains the best PSNR, SSIM and LPIPS within the same runtime, over 0.3 dB above both Adam and LMRS, at half of LMRS's memory. Both the L2 loss and the 16-view batch are out of distribution for our model. 3DGS-LM runs out of memory from the SfM initialization on indoor scenes, and warm-started from 2000 Adam iterations it fits only 13 refinement steps, ending 0.3 dB below Adam at equal time.
| Optimizer | LPIPSββ | SSIMββ | PSNRββ | Iters | Memβ[GB]ββ |
| Second-order | |||||
| LMRS | 0.443 | 0.677 | 25.43 | 335 | 14.2 |
| 3DGS-LM SfM init | OOM on 4/9 scenes | ||||
| 3DGS-LM Adam-2k init | 0.462 | 0.664 | 25.09 | 2k+13 | 33.9 |
| First-order | |||||
| Adam | 0.447 | 0.674 | 25.41 | 4000 | 6.3 |
| L2S-SH0 | 0.432 | 0.682 | 25.77 | 2000 | 6.5 |
The main comparison disables adaptive density control. With the 3DGS-MCMC densification scheme on Mip-NeRF360, both optimizers benefit, and Learn2Splat stays ahead of Adam while gaining more from densification.
| PSNRββ | |||
| MCMC | #G | Adam | L2S |
| w/o | 115k | 25.02 | 25.80 |
| w/ | 178k | 25.12+0.10 | 25.99+0.19 |
Replacing ReSplat's per-pixel error inputs with G3R or Adam gradients raises the peak from 26.6 to 30.0 dB but does not prevent degradation. Adding our meta-training components one at a time on top of the Adam-gradient variant keeps the peak and further reduces the degradation: checkpoint buffer and optimizer rollout stay within 0.7 dB of the peak, the predicted scale factors within 0.13 dB, and our losses remove it entirely, with PSNR still rising at 2000 iterations.
The results above use a single model meta-trained jointly on the sparse and dense configurations, alternating between the two across meta-iterations. Against two dedicated checkpoints, each meta-trained on one configuration only, the unified model is within 0.2 dB on every benchmark. Applying a dedicated checkpoint outside its training regime degrades quality, though the two directions are not symmetric: the dense one loses 0.13–0.36 dB on the sparse benchmarks, while the sparse one loses 7.4–11.9 dB on the dense ones.
| L2SS | L2SD | L2S | ||
| Dense | DL3DV | 17.03 | 28.92* | 29.07 |
| DTU | 21.61 | 29.10* | 29.26 | |
| LLFF | 14.21 | 24.33* | 24.46 | |
| Mip360 | 18.26 | 25.67* | 25.78 | |
| Sparse | DL3DV 8, low-res | 30.60* | 30.24 | 30.53 |
| DL3DV 8, hi-res | 27.10* | 26.90 | 27.17 | |
| DL3DV 32, low-res | 30.18* | 30.05 | 30.25 | |
| RE10K 8 | 26.28* | 26.11 | 26.08 |
Learning rates grid-searched on sparse DL3DV improve every sparse benchmark but degrade most dense ones, so no single Adam configuration is best across regimes and the main results use default Adam. Learn2Splat outperforms default Adam everywhere, and the tuned Adam on all but DTU.
| Adam | Adamtuned | L2S | ||
| Dense | DL3DV | 27.78 | 27.18 | 29.04 |
| DTU | 28.96 | 29.31 | 29.25 | |
| LLFF | 22.82 | 20.98 | 24.56 | |
| Mip360 | 25.62 | 24.42 | 25.75 | |
| Sparse | DL3DV 8, low-res | 30.02 | 30.29 | 30.53 |
| DL3DV 8, hi-res | 26.60 | 26.72 | 27.16 | |
| DL3DV 32, low-res | 29.78 | 30.02 | 30.26 | |
| RE10K 8 | 25.57 | 26.02 | 26.08 |
Learn2Splat reaches higher PSNR early and keeps improving to the end, while G3R⋆ and ReSplat degrade.
*ReSplat is absent from the dense scenes: it relies on pixel-aligned Gaussians from the input views, so it cannot start from SfM points.
If you find this work useful, please cite:
@article{pearl2026learn2splat,
title = {Learn2Splat: Extending the Horizon of Learned 3DGS Optimization},
author = {Pearl, Naama and Esposito, Stefano and Xu, Haofei and Peleg, Amit and Gschossmann, Patricia and
Porzi, Lorenzo and Kontschieder, Peter and Pons-Moll, Gerard and Geiger, Andreas},
journal = {arXiv preprint arXiv:2605.15760},
year = {2026}
}