Skip to main content

One post tagged with "VPU"

VPU stands for Vision Processing Unit, a specialized hardware accelerator designed to efficiently process computer vision and image processing tasks. VPUs are optimized for low-power, high-performance operations, making them ideal for applications in robotics, augmented reality, and autonomous vehicles.

View All Tags

Preloading Multiple ExoPlayer Instances Without Exhausting the VPU's MPS Budget

· 21 min read
Austin Kim
Staff Software Engineer, System programming and media framework specialist

TL;DR

  • The problem: preloading several ExoPlayer instances so playback can start instantly hits the VPU's shared decoder admission budget — creating too many MediaCodec hardware decoders concurrently gets rejected (OMX_ErrorInsufficientResources / MediaCodec.CodecException, ExoPlayer error codes 4001/4003) — even though only one video is ever actually playing at a time.
  • Why it happens: ExoPlayer sets KEY_OPERATING_RATE to the real content frame rate — effectively "1x," the actual playback rate — on every decoder it creates, whether that decoder is actively playing or just sitting preloaded and idle (Section 5), and it never sets MediaFormat.KEY_PRIORITY at all (confirmed absent from MediaCodecRenderer.java/MediaCodecVideoRenderer.java). The net effect: the VPU's admission accounting has no signal telling it "this decoder isn't playing yet" and treats every created instance as if it needs guaranteed real-time decode throughput the moment it's created, not just once it's actually feeding frames. (Qualcomm's own kernel driver control technically defaults KEY_PRIORITY to non-realtime — Section 3.5 — but the resource exhaustion this document exists to solve is consistent with the vendor's closed-source Codec2/OMX HAL layer not preserving that default in practice; unverified from here, flagged as a caveat throughout.)
  • The fix: explicitly set KEY_PRIORITY = 1 (non-realtime) and a low KEY_OPERATING_RATE on every decoder in the preload pool while it's idle, then flip both to realtime / full rate only on the one instance that's actually about to play (Section 6). Non-realtime load is charged at roughly resolution × 1 instead of resolution × fps in the admission math (Sections 3.1–3.3), and non-realtime instances are excluded entirely from the per-core clock-fit check that gates new decoder admission (Section 3.6). For 30fps video, that's roughly a 30× difference in admission cost between an idle non-realtime decoder and one actively playing at realtime priority (worked example, Section 3.4) — meaning, in principle, on the order of 30× more preloaded decoder instances can coexist in the same hardware budget than the naive "every decoder is realtime by default" approach allows.