Day 4 Friday 26 June 2026

The final day of the conference focused on the practical application of Simulation-Based Inference (SBI) and deep learning to bridge the gap between complex simulations and observational data. Presentations spanned a wide range of scales, from the internal properties of stars and the circumgalactic medium to the large-scale cosmic web and galaxy merger histories. A recurring theme was the pursuit of computational efficiency, with several talks demonstrating how neural posterior estimators and differentiable simulations can replace slow traditional sampling methods. The day concluded with a look at the broader infrastructure of AI, discussing self-hosted models for education and ambient AI for scientific synthesis.

Accelerating Bayesian inference via SBI and Neural Posterior EstimationQuantifying the role of environment vs. assembly history in galaxy evolutionBridging simulated mock catalogs with photometric and spectroscopic observationsDifferentiable programming and high-resolution zoom-in simulationsLocal and private AI deployment for academic and professional use
24 slides

Connecting Environment and Assembly to Measure the Formation Histories of Dark Matter Halos With GNN-powered SBI

2026-06-26T10:06:44
Christian Kragh Jespersen — Princeton University

The talk explores the relationship between the spatial environment of dark matter halos and their temporal assembly histories using Graph Neural Networks (GNNs). The speaker demonstrates that GNNs can emulate galaxy properties using environmental data as effectively as they can using merger trees, suggesting a fundamental equivalence between environment and assembly history. This equivalence enables the use of Simulation-Based Inference (SBI) to reconstruct the mass accretion histories of halos from observed cosmic web structures.

  • Galaxy properties are traditionally modeled as functions of either dark matter assembly history (via merger trees) or spatial environment.
  • GNNs were used to emulate galaxy properties (e.g., stellar mass, gas mass) by treating the halo environment as a graph with a tunable linking length.
  • A 'break length' in prediction accuracy was identified, which varies depending on the timescale of the galaxy property being predicted.
  • The results show that environmental information encoded in a GNN can reach the same predictive performance as explicit assembly histories, even in semi-analytic models (SAMs) that assume environment doesn't matter.
  • This equivalence allows for the inference of halo mass accretion histories (e.g., time to reach 25%, 50%, and 75% of current mass) using only environmental data.
  • The method was validated across different simulation types, including SAMs and the TNG magnetohydrodynamic simulations.
Dark Matter HalosGraph Neural NetworksSimulation-Based InferenceGalaxy EvolutionCosmic Web
Abstract (conference schedule)

Modelling the connection between galaxies and their host dark matter halos is a fundamental task in galaxy evolution. The modern use of Graph Neural Networks has allowed a deeper exploration of these relationships on both a galaxy-by-galaxy and population level, demonstrating that galaxies, halos, and their relationships are greatly influenced by both their detailed spatial environments and temporal assembly history. In this talk, I demonstrate a full equivalence in the impact of halo environments and assembly histories on a broad range of baryonic galaxy properties. This result holds when simulating galaxy properties using both full magnetohydrodynamic codes or semi-analytic models, demonstrating that the equivalence is driven by the effect of dark matter on galaxies. To achieve the equivalence, the environment cannot be expressed as spherically averaged density, but must be encoded as a graph, linking halos on sufficiently large spatial scales. We furthermore measure the linking lengths that give optimal predictions for each galaxy property, and show that these are directly related to the typical extent of all halo progenitors. The equivalence offers a tantalizing opportunities for Geometric Deep Learning Simulation-Based inference of the assembly histories of host halos in the real Universe. I will finish by showing some early results demonstrating our GNN-SBI’s capability to directly infer key events in the formation of dark matter halos.

Questions for the speaker (3)
  1. How does the choice of linking length for the GNN graph specifically relate to the physical 'turnaround radius' or the typical extent of halo progenitors?
  2. Given that SAMs explicitly assume galaxy properties are a function of assembly history alone, why does the GNN find an equivalent mapping using only the current spatial environment?
  3. What are the primary sources of uncertainty or degeneracy when using SBI to backtrack the mass accretion history of a halo from its current environmental graph?
10 slides

Introducing L-Galaxies AD: an auto-differentiable galaxy evolution simulation

2026-06-26T10:42:02
Andrew Green — University of Hertfordshire

The speaker presents L-Galaxies AD, a complete rebuild of the L-Galaxies semi-analytic model (SAM) in C++ to enable automatic differentiation (AD) and process-level parallelism. By transforming the simulation into a differentiable computational graph, the framework allows for the use of gradient-based Bayesian inference methods like Hamiltonian Monte Carlo (HMC) to calibrate high-dimensional parameter spaces more efficiently than traditional MCMC.

  • Rebuilt the legacy C-based L-Galaxies code into a modular C++ framework to improve maintainability and testability.
  • Implemented automatic differentiation using operator overloading and templates via the CoDiPack library.
  • Utilized a computational graph representation to identify independent branches for parallelism and to facilitate reverse-mode AD (backpropagation).
  • Enabled the use of gradient-aware samplers such as HMC and NUTS to overcome the scaling bottlenecks of random-walk MCMC in a ~15-parameter space.
  • Planned applications include calibrating dwarf-galaxy physics against the EDGE simulation and integrating radio continuum emission and black hole physics.
Semi-analytic models (SAMs)Automatic DifferentiationBayesian InferenceGalaxy EvolutionComputational Graphs
Abstract (conference schedule)

Semi-analytic models (SAMs), such as L-Galaxies, play a vital role in modelling galaxy populations across cosmological volumes. These models evolve baryonic physics on pre-computed dark matter halo merger trees and achieve much lower computational cost than full hydrodynamical simulations. This efficiency supports robust parameter inference via methods such as Markov Chain Monte Carlo (MCMC), enabling extensive model calibration and comparison with observations. However, as SAM physics become more complex and parameter spaces expand, run times increase, making traditional inference increasingly challenging. We present a new optimised, parallel C++ implementation of L-Galaxies (Yates et al., 2024) that we have made fully differentiable. We ensure differentiability by using operator overloading for forward- and reverse-mode automatic differentiation (AD). By structuring the SAM as a computational graph, we enable process-level parallelism and enable efficient reverse-mode AD. By exposing exact gradients via AD, our framework enables gradient-based inference methods such as Hamiltonian Monte Carlo (HMC), which deliver higher sampling efficiency and faster convergence in complex, high-dimensional parameter spaces. Our differentiable, parallel SAM will be used to generate extensive galaxy catalogues efficiently, and to directly constrain model parameters using Bayesian inference against super-high-resolution hydrodynamical simulations — demonstrating a scalable gradient-aware approach to simulation-based inference for galaxy evolution.

Questions for the speaker (3)
  1. Given that reverse-mode AD can incur a significant memory and computational overhead compared to forward propagation, how does the tape-based approach in CoDiPack impact the memory footprint when simulating large cosmological volumes?
  2. You mentioned that parallel boundaries require extra work for derivative propagation; specifically, how are gradients handled across the synchronization points where independent branches of the computational graph converge?
  3. For the planned calibration against the EDGE hydrodynamical simulations, how do you intend to define the likelihood function or loss metric to compare the SAM's outputs with the high-resolution hydro data?
16 slides

Exploring the Impact of Dust on Galaxy Luminosity Functions with Semi-Analytic Models

2026-06-26T10:55:58
Sophie Newman — University of Portsmouth

The speaker presents a forward-modelling pipeline using the Synthesizer tool to apply flexible, physically motivated dust attenuation to galaxies generated by the SC-SAM semi-analytic model. The work demonstrates that varying dust prescriptions significantly impact luminosity functions in the UV spectrum compared to fixed models like Calzetti. Preliminary results using Simulation-Based Inference (SBI) suggest that colour distributions may be more effective than luminosity functions for recovering dust model parameters.

  • Semi-analytic models (SAMs) allow for rapid generation of large galaxy populations but often lack the spatial data required for full radiative transfer.
  • Fixed dust attenuation curves (e.g., Calzetti, Milky Way) fail to capture the systematic variation in dust properties observed across different galaxy types.
  • The proposed pipeline uses Synthesizer to apply a per-galaxy attenuation model based on metallicity, gas mass, and a calibrated disk radius correction factor.
  • Comparison of luminosity functions shows that the flexible model diverges from the Calzetti model primarily in the FUV and NUV bands, aligning better with some observations.
  • Early tests with the sbi package indicate successful recovery of the radius correction factor, though other parameters like slope and bump strength remain difficult to constrain.
  • Colour distributions currently show better parameter recovery than luminosity functions alone.
Galaxy Luminosity FunctionsSemi-Analytic Models (SAMs)Dust AttenuationForward ModellingSimulation-Based Inference (SBI)
Abstract (conference schedule)

We generate synthetic observations from the SC-SAM semi-analytic model to directly connect galaxy formation physics with observable galaxy properties. Using Synthesizer, we develop a flexible, physically motivated forward-modelling pipeline that incorporates a suite of galaxy dust prescriptions spanning a range of grain-size distributions, dust masses, and compositions. We demonstrate that this flexible dust modelling produces colours and luminosity functions that differ from those obtained using fixed dust models, with the largest discrepancies occurring in the UV. By embedding this framework within a SBI pipeline, observational data could be used to constrain dust physics. Future work will extend the framework to include self-consistent modelling of AGN emission and dust attenuation.

Questions for the speaker (3)
  1. What specific physical justification or empirical data was used to determine the correction factor for scaling the disk radius to the gas radius?
  2. Given that the FUV and NUV luminosity functions show different levels of agreement with observations, which specific dust parameters (e.g., bump strength vs. slope) are driving these differences?
  3. Why do colour distributions provide better parameter recovery in your SBI tests compared to luminosity functions, and would combining both observables significantly improve the constraints on slope and intercept?
24 slides

Quantifying Environmental Information in Galaxy Evolution with Interpretable Machine Learning

2026-06-26T11:14:29
Shun-ya S. Uchida — Nagoya University

The speaker presents an interpretable machine learning framework to quantify how galaxy properties are influenced by both their local spatial environment and primordial initial conditions. Using hydrodynamical simulations, the work employs SHAP analysis and convolutional neural networks to disentangle the effects of halo mass from secondary environmental drivers across different galaxy types.

  • Developed a neural network to predict stellar mass and star formation rate (SFR) using both host halo properties and the properties of the nearest 30 neighboring halos.
  • Used SHAP values to quantify the 'environmental contribution' ratio, finding that low-mass and satellite galaxies are more strongly influenced by their environment than high-mass central galaxies.
  • Implemented a 3D-CNN to map initial density and velocity fields at z=20 to z=0 galaxy properties, including halo mass as an auxiliary output to improve learning.
  • Discovered that low-mass galaxies require predictive information from a larger spatial region beyond the critical radius compared to high-mass galaxies.
  • Identified a negative correlation between initial velocity dispersion and final stellar mass, suggesting that lower dispersion facilitates more efficient gravitational collapse.
  • Proposed future work using Simulation-Based Inference (SBI) and Normalizing Flows to model the distribution of quenched galaxies.
Galaxy EvolutionInterpretable Machine LearningSHAP AnalysisCosmological SimulationsInitial Conditions
Abstract (conference schedule)

Understanding galaxy evolution requires disentangling the nonlinear connections between galaxies, dark matter, and their environments across cosmic time. We present an interpretable machine-learning framework that quantifies two complementary definitions of environment. First, we define the spatial environment as the distribution of neighbouring galaxies at z = 0. Using hydrodynamical simulations, we train a neural network to predict galaxy properties from host and neighbouring subhalo properties. Explainable AI analysis then quantifies environmental parameters driving secondary dependencies beyond halo mass, as a function of galaxy type and mass, — providing physical insight into the environmental effect of galaxy formation. Second, we define the primordial environment as the initial dark matter field that causally encodes present-day galaxy properties through structure formation. We train a convolutional neural network that maps the initial density and velocity fields at z = 20 directly to z = 0 galaxy properties. Explainable AI analysis reveals which spatial regions and patterns in the initial conditions carry predictive information for baryonic observables, extending recent CNN-based halo-mass prediction studies to galaxy properties. This forward-modelling approach — from initial conditions to galaxy observables — is relevant to field-level inference with SBI. Our results offer quantitative, interpretable insights into the nonlinear encoding of initial conditions in galaxy observables.

Questions for the speaker (3)
  1. How does the use of a fixed-size input box for initial conditions at z=20 bias the predictions for high-mass halos that may have originated from regions larger than the box?
  2. You mentioned that the environmental contribution is highest for satellite galaxies; to what extent does this result depend on the specific choice of the 30 nearest neighbors versus a fixed spatial radius?
  3. Regarding the negative correlation between initial velocity dispersion and stellar mass, does this relationship hold across different galaxy types, or is it primarily driven by the most massive systems?
5 slides

Fast Stellar Age Inference for Galactic Archaeology with SBI

2026-06-26T12:01:18
Shichen Su — UCL

The speaker presents a pipeline for determining stellar ages using Simulation-Based Inference (SBI) to overcome the computational bottlenecks of traditional Bayesian isochrone fitting. By training a Neural Posterior Estimator (NPE) on MIST isochrones and implementing a survey-agnostic noise handling method at inference time, the approach achieves speeds 10x faster than nested sampling with comparable accuracy.

  • Stellar age determination is critical for galactic archaeology but computationally expensive using traditional iterative sampling for millions of stars.
  • The pipeline uses a Neural Posterior Estimator (NPE) trained on MIST isochrones, utilizing Equivalent Evolutionary Points (EEP) as a proxy for mass to improve interpolation.
  • To remain survey-agnostic, noise is handled at inference time by marginalizing over multiple noise-free realizations of the observed data.
  • Validation against nested sampling using medium-resolution noise shows similar accuracy, with a mean absolute error of 0.2 dex for age.
  • The method is well-calibrated for turn-off stars, though it shows some over-confidence for main-sequence stars where observables are less sensitive to age.
  • The SBI approach is approximately 10 times faster than nested sampling, making it scalable for large surveys like DESI, Gaia XP, and 4MOST.
Simulation-Based InferenceStellar Age InferenceGalactic ArchaeologyNeural Posterior EstimationIsochrone Fitting
Abstract (conference schedule)

Understanding the assembly history of the Milky Way requires precise stellar ages across large populations. These ages help to disentangle different merger events and constrain galactic chronology. Large-scale spectroscopic surveys are central to this effort: DESI and Gaia XP have already delivered stellar parameters (e.g., Teff and [Fe/H]) for millions of stars, and upcoming surveys like WEAVE and 4MOST will expand this dataset further. These parameters break degeneracies around the main-sequence turn-off and subgiant branch, enabling precise age-determination via isochrone fitting in these regions. However, traditional Bayesian methods require running iterative sampling algorithms independently for each star, creating a computational bottleneck when scaling to millions of targets. Simulation-based inference offers a solution to this problem through amortised inference. In this work, we train a neural posterior estimator on MIST isochrones that can be deployed across different surveys with varying noise levels, bypassing the per-star computational cost of conventional methods. We present preliminary validation results to demonstrate the feasibility of population-scale age inference across current and future spectroscopic surveys.

Questions for the speaker (3)
  1. How does the pipeline handle cases where a noisy observation falls outside the training distribution of the NPE, and how would the proposed probability tracking specifically identify these unreliable estimates?
  2. Given that the model is slightly over-confident for main-sequence stars, what specific modifications to the training data or the noise marginalization process could improve calibration in that evolutionary stage?
  3. How will the integration of parallax data into the current NPE framework change the dimensionality of the input space and the complexity of the noise-free realization sampling?
18 slides

Simulation Based Inference for Understanding the Role of AGN in Galaxy Evolution

2026-06-26T12:30:44
Thomas Bebbington — University of Bristol

The talk proposes using Simulation Based Inference (SBI) to accelerate the extraction of Active Galactic Nuclei (AGN) properties from large-scale surveys like Euclid. The speaker discusses the challenges of using cosmological hydrodynamical simulations as priors, specifically how varying sub-grid physics implementations across different simulations lead to divergent results.

  • Traditional SED fitting is too slow for the volume of data expected from the Euclid survey.
  • SBI can provide a rapid inference pipeline by learning the mapping between simulation parameters and synthetic observables.
  • Cosmological simulations (e.g., Eagle, TNG, Simba) differ significantly in their black hole accretion and feedback implementations, such as seed mass and boost factors.
  • Sub-grid physics parameters are often numerical and resolution-dependent rather than directly physically meaningful.
  • Proposed research directions include testing how different observational calibration data affect inferred parameters and performing Bayesian model comparisons between different simulation physics.
  • Comparison of TNG and Simba populations shows that different modeling choices create distinct, sometimes bimodal, distributions in properties like dust mass.
Simulation Based InferenceActive Galactic NucleiGalaxy EvolutionCosmological SimulationsEuclid Survey
Abstract (conference schedule)

Euclid will observe close to one billion galaxies across a wide range of cosmic time, providing opportunity to investigate the processes behind galaxy evolution and the influence Active Galactic Nuclei (AGN) have on their host systems. Maximising the scientific insights from such large datasets requires inference frameworks that can quickly and efficiently connect observations to theoretical models. AGN are known to play an important role in suppressing galaxy and stellar formation, providing feedback to their local surroundings. In order to understand their role and impact in galaxy evolution, detailed modelling is important to capture all the physics involved. To reflect the data gathered by Euclid, this modelling needs to take place within cosmological scale hydrodynamical simulations. However, at these scales numerical resolution limitations prevent AGN from being resolved directly, requiring their feedback to be implemented through sub-grid models. The parameterisation and calibration of these models strongly affect simulated galaxy properties, leading to modelling uncertainties that limit their comparison with survey data. By combining cosmological simulations and the software package Synthesizer to generate realistic synthetic observables, I plan to use simulation based inference and Euclid data to directly compare cosmological simulation predictions to observational data, with a particular focus on the AGN sub-grid physics implemented. I will outline a general framework for applying SBI techniques to large extragalactic surveys such as Euclid, highlighting their potential to advance our understanding of the role AGN play in galaxy evolution and cosmological simulations.

Questions for the speaker (3)
  1. Given that sub-grid parameters are often resolution-dependent and numerical, how do you plan to map these inferred values back to physically meaningful AGN properties?
  2. You mentioned that different simulations like TNG and Simba produce distinct populations; how will you handle the potential bias introduced by choosing one specific simulation as the prior for your SBI framework?
  3. In your proposed model comparison, what specific metrics or Bayesian evidence criteria will you use to determine if one feedback implementation is more 'physically motivated' than another?
30 slides

Deep learning approaches for galaxy merger identification and classification: bridging simulations and observations

2026-06-26T12:49:40
Subhrata Dey — National Centre for Nuclear Research

The speaker presents a supervised deep learning framework to classify galaxies into non-merger, pre-merger, and post-merger stages using an ensemble of convolutional neural networks. The models were trained on mock IllustrisTNG images forward-modeled to match Hyper Suprime-Cam (HSC) observations and subsequently applied to the North Ecliptic Pole (NEP) field to create a science-ready merger catalog.

  • Developed a three-class classifier (non-merger, pre-merger, post-merger) to capture distinct physical properties across merger stages.
  • Utilized an ensemble of eight CNN backbones, incorporating different architectures and image scaling methods (arcsinh and log-normal) to improve accuracy and precision.
  • Demonstrated that ensemble agreement (e.g., 8/8 models agreeing) significantly increases classification confidence and precision, particularly for difficult-to-detect post-mergers.
  • Analyzed performance trends showing that non-merger precision decreases at higher stellar masses, likely due to high-mass galaxies exhibiting merger-like morphologies.
  • Applied the framework to the AKARI-NEP field, identifying thousands of pre- and post-merger galaxies and analyzing their distribution relative to local environment density.
  • Found a positive correlation between merger fraction and local density, though post-mergers showed no such increasing trend.
Galaxy MergersDeep LearningConvolutional Neural NetworksIllustrisTNGHyper Suprime-Cam (HSC)Ensemble Learning
Abstract (conference schedule)

Galaxy mergers play a central role in the hierarchical assembly of galaxies, driving bursts of star formation, triggering AGN activity, and reshaping galaxy structure. However, robustly identifying mergers and distinguishing their evolutionary stages (pre-merger and post-merger) from imaging data alone remains challenging, particularly in the era of wide-field surveys with millions of galaxies. We present a supervised deep learning framework trained on mock IllustrisTNG galaxy images forward-modeled to match Hyper Suprime-Cam (HSC) observations (Margalef-Bentabol et al. 2024). Using convolutional neural networks (CNNs), we develop a three-class classifier (non-merger, pre-merger, post-merger). To enhance robustness and mitigate architecture-dependent biases, we train multiple CNN backbones and combine them within an ensemble framework. This ensemble strategy leverages complementary feature sensitivities across models, significantly improving classification accuracy and reliability. We evaluate performance on both synthetic and real HSC data, assessing domain transfer and generalizability, key challenges for simulation-based inference in galaxy evolution. The resulting framework produces science-ready merger catalogs for the North Ecliptic Pole (NEP) field, enabling statistically robust studies of merger-driven star formation, black hole growth, and structural transformation.

Questions for the speaker (3)
  1. Given that non-merger precision drops at higher stellar masses, how do you distinguish between genuine merger features and the intrinsic structural disturbances common in massive elliptical galaxies?
  2. You noted that the merger fraction decreases at higher redshifts due to observational limits; have you quantified the specific surface brightness or resolution threshold where post-merger features become undetectable?
  3. Why did the post-merger fraction show no increase with local density, while pre-mergers did, and what does this imply about the timescales of the post-merger phase in dense environments?
24 slides

Inferring CGM Cloud Properties Using Simulation-Based Inference

2026-06-26T15:19:06
Tanmay Singh — Arizona State University

The speaker presents a framework to infer the physical properties of the circumgalactic medium (CGM) by applying simulation-based inference (SBI) to synthetic absorption spectra. By leveraging the SALSA catalog from cosmological simulations, the work aims to overcome the degeneracies of traditional Voigt-profile fitting and recover latent gas properties and origins.

  • Traditional Voigt-profile fitting often fails to recover weak absorbers and suffers from degeneracies in gas configuration.
  • Comparison of SIMBA and TNG50 simulations shows that both underpredict the OVI covering fraction compared to observations, partly due to the difficulty of detecting weak absorbers.
  • The proposed SBI pipeline uses a 1D-CNN embedding and Neural Spline Flows to invert the forward model and estimate posteriors for cloud and halo properties.
  • The framework aims to infer non-observable latent variables, such as whether gas originates from accretion or satellite galaxy stripping.
  • Preliminary application involves matching observed high-velocity clouds (e.g., in M61) with similar structures in TNG50 to backtrack their origins.
Circumgalactic Medium (CGM)Simulation-Based Inference (SBI)Absorption SpectroscopyCosmological SimulationsNeural Spline Flows
Abstract (conference schedule)

Understanding the physical state of the circumgalactic medium (CGM) is central to galaxy evolution, but we can only primarily probe it indirectly through absorption lines in background quasar spectra. The standard approach of fitting Voigt profiles and running ionization models works, but it is fundamentally degenerate: very different gas configurations can produce nearly identical absorption features, blended components are hard to disentangle, and propagating uncertainties through the full pipeline is difficult. In this work, we develop a SBI framework that learns to infer multiphase CGM gas properties directly from spectra, trained entirely on synthetic data from the SALSA catalog. SALSA provides millions of sightlines through IllustrisTNG, SIMBA, and EAGLE, each with a realistic synthetic spectrum and full knowledge of every gas cell the sightline intersects. We group these gas cells into physically coherent clouds by clustering in velocity, density, and temperature space, giving us supervised labels for each spectrum: the number of clouds, their individual properties (density, temperature, metallicity, column density, velocity, path length), and halo-scale environmental quantities. Our model uses neural posterior estimation with a 1D convolutional or Transformer encoder feeding into conditional normalizing flows, with physics-informed priors on thermal broadening, density–temperature relations and doublet ratios. We plan to validate first on SALSA sightlines and then apply the framework to real HST/COS data from CGM surveys, benchmarking against traditional Voigt-profile and photoionization analyses. This extends SBI into the structured, multiphase CGM regime.

Questions for the speaker (3)
  1. Given that both SIMBA and TNG50 underpredict the OVI covering fraction, how does this systematic discrepancy in the training data affect the accuracy of the SBI posteriors when applied to real HST/COS data?
  2. What specific architectural choices were made for the 1D-CNN embedding to ensure it captures the necessary spectral features for distinguishing between clumpy and diffuse gas configurations?
  3. How do you plan to quantify the uncertainty or reliability of the inferred 'latent observables,' such as the origin of the gas, since these cannot be directly verified with observational data?
26 slides

Disentangling Feedback and Variance in 1,024 Milky Way-Mass DREAMS Simulations

2026-06-26T15:19:06
Jonah Rose — University of Florida

The speaker presents the DREAMS project, a suite of 1,024 high-resolution cosmological zoom-in simulations designed to separate the effects of sub-grid physics from intrinsic halo-to-halo variance. Using a simulation-based inference framework and a pseudo-posterior weighting scheme, the talk demonstrates how to identify realistic parameter spaces and quantify the dominance of cosmic variance over model uncertainty in galaxy properties.

  • DREAMS utilizes high-resolution zoom-in simulations of Milky Way-mass halos to resolve satellite galaxies and internal structures better than uniform box simulations.
  • A pseudo-posterior inference pipeline was implemented using the SAGA stellar mass-halo mass relation as a baseline summary statistic to weight the parameter space.
  • The inference reveals broad degeneracies in feedback parameters and shows that the high-likelihood region is offset from fiducial TNG parameters due to resolution differences.
  • Results indicate that while sub-grid physics affects galaxy mass, intrinsic cosmic variance dominates population statistics for satellites.
  • Specific merger histories, such as Gaia-Sausage-Enceladus-like events, correlate with quenched galaxies and smaller discs, though significant variance remains.
  • The project is expanding to include a wider range of halo masses (10^9 to 10^14) to recreate full-box statistics using a suite of zoom-in simulations.
Cosmological SimulationsSimulation-Based InferenceGalaxy FormationSub-grid PhysicsCosmic Variance
Abstract (conference schedule)

We introduce a simulation-based inference framework built upon the DREAMS Project, a new suite of 1,024 cosmological hydrodynamical zoom-in simulations of Milky Way-mass halos [2512.00148]. Designed to systematically disentangle theoretical uncertainties in sub-grid physics from intrinsic halo-to-halo variance, this suite varies key astrophysical parameters governing supernova wind energy, wind speed, and AGN feedback efficiency within the IllustrisTNG model. We introduce a novel observational weighting scheme, applying pseudo-posterior constraints based on the empirical stellar mass-halo mass relation [2602.03613]. This inference pipeline reveals broad degeneracies in the fiducial feedback parameters, demonstrating that standard single-model tuning misses complex parameter interdependencies. Using these constraints, we quantify that a Gaia-Sausage-Enceladus-like merger history is exceptionally rare and drives specific structural shifts, though immense halo-to-halo scatter persists. For satellites, our emulator demonstrates that while sub-grid variations regulate stellar mass, intrinsic variance overwhelmingly dominates population statistics [2512.02095]. We also identify a persistent tension where the inferred parameter space fails to reproduce the extended half-light radii of massive satellites observed in the SAGA survey. To determine if this size-mass tension is an artifact of the TNG sub-grid model or a broader theoretical challenge, we will also present preliminary results from ongoing extensions to the DREAMS project framework. This includes three new suites of high-resolution dwarf galaxy simulations run with FIRE3, RAMSES, and ChaNGa. Furthermore, we will introduce new mass-varied and resolution-varied simulation suites utilizing the IllustrisTNG and FIRE3 models. These suites extend the halo mass range to include masses from 10^9 to 10^14, allowing us to recreate uniform box results from a suite of zoom-in simulations.

Questions for the speaker (3)
  1. How does the choice of the SAGA stellar mass-halo mass relation as the sole summary statistic impact the resulting pseudo-posteriors compared to a full field-level inference?
  2. You mentioned that the high-likelihood parameter region is offset from the fiducial TNG parameters due to resolution; can you quantify the magnitude of this resolution dependence?
  3. Regarding the recreation of full-box statistics from zoom-ins, how do you account for the potential bias introduced by selecting halos based purely on mass without considering environmental density?
20 slides

Image-Based Inference: Bridging Photometric Surveys and Spectroscopic Constraints

2026-06-26T15:57:40
James Kostas Ray — UCL

The speaker presents a likelihood-free inference framework that uses pixel-level imaging data to predict the posterior distributions of galaxy physical properties, bypassing the need for simulated intermediaries. By leveraging an ensemble of CNNs and Vision Transformers (ViT), the method maps morphologically rich photometric data to high-fidelity spectroscopic labels to enable scalable population-level inference.

  • Addresses the bottleneck where photometric surveys observe billions of galaxies but spectroscopic surveys target far fewer.
  • Utilizes galaxy morphology as a primary constraint to improve the accuracy of physical parameter extraction.
  • Employs an ensemble of CNNs and Vision Transformers (ViT) to capture different inductive biases in the latent space.
  • Demonstrates high computational efficiency, capable of generating a billion galaxy posteriors in approximately three months on a consumer GPU.
  • Uses neural density estimation to capture degenerate posteriors, providing a measure of model uncertainty.
  • Briefly discusses the application of reinforcement learning (RL) for telescope scheduling (Ariel) and optimizing galaxy redshift tomographic bins.
Likelihood-free inferenceGalaxy morphologyNeural density estimationPhotometric redshiftsReinforcement learning
Abstract (conference schedule)

Modern galaxy surveys are increasingly photometry dominated with wide field imaging campaigns observing orders of magnitude more objects than spectroscopic surveys can feasibly target. While spectroscopy provides physically interpretable constraints on stellar populations, star formation histories, metallicities, and dust, its limited coverage creates a bottleneck for population level inference. This imbalance demands new methodologies that transfer spectroscopic information into the vastly larger imaging domain. This work develops a likelihood free, image based inference framework that directly maps observed multiband galaxy images to posterior distributions over spectroscopically derived physical properties. Only recently have hydrodynamical simulations begun to produce galaxy morphologies that resemble those observed in modern surveys. Rather than relying on simulated images as an intermediary, we instead learn the mapping directly from real data. The approach operates on pixel level imaging information, allowing morphology and structural features to act as primary constraints on physical properties.

Questions for the speaker (3)
  1. How does the model handle the 'implicit prior' and the risk of inheriting systematics from the spectroscopic training labels when mapping to the larger photometric dataset?
  2. What specific morphological features are the CNN and ViT architectures capturing differently in the latent space to justify the use of an ensemble?
  3. In the reinforcement learning approach for redshift optimization, how is the reward function defined to maximize the angular cross-correlation signal between photometric and spectroscopic samples?
6 slides

Self-Hosted AI for Education

2026-06-26T17:04:11
Sotiria Fotopoulou

The speaker discusses the implementation of a self-hosted AI tutor for physics students using Open WebUI and the Gemma model. She compares this approach to fine-tuning and proprietary tools, emphasizing the importance of equitable access and data privacy in an academic setting.

  • Motivation for self-hosting includes ensuring equitable access for students, maintaining privacy of internal documents, and avoiding token costs of paid services.
  • Comparison of AI implementation strategies: fine-tuning a small model (laborious and hard to scale), using Google's NotebookLM (not open source/university-provided), and self-hosting with Open WebUI.
  • Demonstration of creating specialized AI personas (Expert, Support, Teacher) by combining a base model (Gemma) with specific system prompts and curated knowledge bases.
  • Technical stack overview involving Docker containers, Ollama for model routing, and optional integrations like PostgreSQL for embeddings and Presidio for data masking.
  • Warning regarding administrative visibility, noting that self-hosted system admins can see all user interactions.
  • Challenges identified include the effort required to curate a high-quality knowledge base and the risk of information becoming stale in rapidly evolving technical fields.
Self-hosted AIOpen WebUIRAG (Retrieval-Augmented Generation)Educational TechnologyLLM Personas
Questions for the speaker (3)
  1. How did you determine the optimal balance of knowledge base content to prevent the model from ignoring its general training data, as you mentioned it sometimes gets 'stuck' in the provided knowledge?
  2. Given that you want to avoid providing students with overly long, 'perfect' answers, what specific prompting techniques or model parameters are you using to tune the tutor for an introductory learning level?
  3. What specific metrics or validation methods do you plan to use to ensure the curated knowledge base is accurate and that the AI tutor is not hallucinating technical instructions for the students?
9 slides

Ambient AI for Scientific Meetings

2026-06-26T17:14:57
Chris Lovell

The speaker presents an ambient AI system deployed during a conference to provide live transcription, automated talk summaries, and a final synthesis white paper. The system emphasizes privacy by running all transcription and LLM processing on local GPU hardware.

  • Implemented a local pipeline using a 13B parameter Gemma model to ensure data privacy and avoid external AI companies.
  • The system captures live audio and screen content to generate per-talk summaries, daily overviews, and suggested questions for speakers.
  • A final conference white paper was generated by synthesizing all transcripts and screenshots, comparing current trends (e.g., diffusion models) against previous years' abstracts.
  • Manual intervention was required to mark the end of talks due to a 30-second stream delay that hindered automated detection.
  • The project was developed by Toby Lovick and is available as a public repository for others to implement at their own meetings.
Ambient AILocal LLMsAutomatic Speech RecognitionPrivacy-Preserving AIConference Synthesis
Questions for the speaker (3)
  1. How did the 13B Gemma model's performance in summarizing technical scientific content compare to the frontier models used for the final white paper?
  2. What specific audio capture hardware or techniques were used to ensure that audience questions were captured clearly enough for accurate transcription?
  3. Beyond the 30-second stream delay, what were the primary failure modes of the automated screen capture and transcription synchronization?