Preprint · not yet peer-reviewed · 2026

Where the Time Goes: Profiling, Segmenting and Verifying CPU and GPU Rendering in a Narrated-Video Pipeline

Rajesh Kumar Kona

Independent researcher · Edinburgh, United Kingdom · 25 September 2026

Download PDF Technical report (long version) Voice-to-Video on the portfolio All papers

Abstract

Automatic explainer-video systems draw every frame themselves, and rendering is the one cost that grows with the length of the video. I study where that time goes in Voice-to-Video, my narration-to-video pipeline, whose reference renderer composes frames with Pillow and encodes them with x264, and I ask how the work should be split between CPU and GPU. Measured per content kind at 1080p, composition takes 54% of CPU time for typography, 74% for mixed content and 85–90% for photographs and transitions; a profile of a real project put it at 77%. By Amdahl's law that caps a hardware encoder at 1.30×, while removing redundant buffer copies alone took composition from 12 to 26 frames/s. Three designs follow: frame-exact segmented rendering, which makes renders resumable and parallel; a narrow OpenGL 3.3 compositor that offloads only photograph resampling; and an equivalence contract that admits a GPU only if no channel of any pixel differs from the CPU reference by more than 2 levels. The two-pass Lanczos-3 shader meets the contract on an NVIDIA GeForce GTX 1650 (23/23 scenes, worst difference 2) and on Mesa llvmpipe (worst 1). Ablations show the clamp between passes, dropped edge taps and the Lanczos kernel are each necessary (without them the worst differences are 23, 57 and 84). On the GTX 1650 at 1080p the GPU composes photographs 11.7× faster than the CPU (3.7 to 43.3 frames/s), transitions 7.7× and mixed content 3.0×; typography is 0.94× and stays on the CPU. Once photographs compose on the GPU, x264 becomes the larger share of the work (64%), so the case for hardware encoding has to be reopened. Two- and four-hour renders finished with no memory growth and zero duration error.

Keywords: video rendering, GPU compositing, image resampling, Lanczos, Amdahl's law, checkpointing, differential testing, x264

Cite

@misc{kona2026wherethetimegoes,
  author = {Kona, Rajesh Kumar},
  title  = {Where the Time Goes: Profiling, Segmenting and Verifying
            {CPU} and {GPU} Rendering in a Narrated-Video Pipeline},
  year   = {2026},
  note   = {Preprint},
  url    = {https://0krk0.dev/papers/voice-to-video-rendering.html}
}

This is a preprint. It has not yet been peer-reviewed. Every number in it comes from the project's own test harnesses or from the experiments described in the paper.