Hardware acceleration
Software encoding scales with CPU count. A high-quality nine-rung encode ladder pushes a 16-vCPU instance to 95% capacity — one channel per machine. NVIDIA NVENC and NETINT Quadra VPU offload encode and decode to dedicated hardware, running the same ladder at a fraction of the CPU load and a fraction of the cost per channel. Norsk detects available hardware and routes tasks to it automatically. No configuration of offload paths required.
How it works
Norsk workflows are hardware-agnostic at the point of authorship. You build the workflow once; Norsk decides at runtime where each task runs.
A nine-rung H.264 encoding ladder on a high-quality profile pushes a 16-vCPU CPU-only instance to 90–95% utilisation. That instance handles one channel. At ten concurrent channels you need ten instances. At hundreds of concurrent live events, CPU cost is the dominant line in the infrastructure budget — and the only lever is more hardware. Upgrading quality settings or adding rungs to the ladder makes it worse. The CPU scaling problem is not a tuning problem; it is a structural one.
NVIDIA NVENC offloads H.264 and HEVC encode and decode to the GPU. NETINT Quadra uses a purpose-built ASIC architecture (Codensity G5) for H.264 and HEVC encode and decode at 8-bit or 10-bit. Both run full encoding ladders at dramatically lower CPU utilisation. The nine-rung Apple TN2224 ladder that saturates a 16-vCPU CPU instance at 95% runs on a four-vCPU NVIDIA T4 instance at 40% CPU and 54% GPU — the same quality, on a smaller instance. The same ladder on a NETINT Quadra T1U runs at 30% CPU and 19% VPU on eight vCPUs, with headroom left on both. Multiple channels fit where one ran before.
Norsk detects the hardware available at runtime and routes encode and decode tasks to it without manual configuration of offload paths. When a processing step — scaling, picture-in-picture composition, frame rate conversion — is natively supported by the hardware, it runs on the card. The NETINT Quadra handles scaling and composition directly on the VPU, keeping the media on hardware end to end. When a step is not supported by the available hardware, Norsk falls back to CPU for that step and returns the result to the card for encode. The same workflow file runs on a CPU-only server, an NVIDIA GPU instance, or a NETINT Quadra machine without modification.
Use cases
Hardware acceleration changes the economics of high-volume encoding. These are the operations where that change is most measurable.
A live event operator running hundreds of concurrent streams finds that the per-channel CPU cost, manageable at ten channels, is a significant operating expense at scale. Increasing encode quality or expanding the ABR ladder requires proportionally more CPU instances. The engineering team is spending time managing instance counts rather than improving the product.
On NETINT Quadra hardware, a 2U server with 24 Quadra cards delivers hundreds of channels simultaneously. On NVIDIA GPU instances in the cloud, the same encoding ladder that required a 16-vCPU CPU instance runs on a 4-vCPU GPU instance with capacity to spare. Per-channel infrastructure cost falls; channel density rises. The engineering team manages fewer machines.
Talk to us about your encoding infrastructure →A broadcast facility owns its infrastructure and wants to maximise channel throughput from existing rack space without adding more servers. CPU-based encoding delivers one channel per high-spec server at broadcast quality. The facility has the space and the power budget for a fixed number of machines; adding more is a capital decision.
NETINT Quadra cards install into standard servers. A single 2U server with 24 Quadra cards runs Norsk workflows across hundreds of channels — encode, decode, scaling, and composition all on the VPU. The existing rack handles the load that would otherwise require a room full of CPU servers. Norsk manages the workflow and hardware routing; the facility manages the server.
Talk to us about on-prem deployment →A streaming operator running cloud encoding benchmarks finds that the instance types recommended for software encoding — 16-vCPU compute-optimised instances — are cost-competitive at low channel counts but become expensive at scale. GPU instance types exist but the engineering team is unsure which instance, which profile, and which ladder configuration produces the best cost-to-quality ratio.
Norsk has benchmarked standard industry encoding ladders — Apple TN2224, Netflix one-size-fits-all, AWS Elemental — across NVIDIA T4 (EC2 g4dn) and NVIDIA RTX 4000 Ada (Akamai/Linode) GPU instances, with CPU utilisation, GPU utilisation, and RAM usage published for ultra-low latency, balanced, and high-quality profiles. The starting point for instance selection is documented rather than derived by trial.
Talk to us about cloud encoding →Capabilities
Both are fully supported in Norsk workflows. The right choice depends on your deployment model, your quality targets, and whether you are in the cloud or on-prem.
NVIDIA NVENC offloads H.264 and HEVC encode and decode to the GPU, running full ABR encoding ladders at a fraction of the CPU load. Benchmarked on NVIDIA T4 (AWS EC2 g4dn) and NVIDIA RTX 4000 Ada (Akamai/Linode). On the nine-rung Apple TN2224 high-quality profile: 40% CPU + 54% GPU on a four-vCPU T4 instance, versus 95% CPU on a 16-vCPU CPU-only instance. Available on AWS and Akamai cloud instances.
NETINT Quadra uses the Codensity G5 ASIC architecture for H.264 and HEVC encode and decode at 8-bit or 10-bit. The Quadra handles scaling, resizing, and picture-in-picture composition natively on the card, keeping media on hardware through the full processing chain. A 2U server with 24 Quadra cards delivers hundreds of concurrent channels. Benchmarked on NETINT Quadra T1U: nine-rung high-quality profile at 30% CPU and 19% VPU on eight vCPUs.
The NETINT T408 VPU is also supported. Encode and decode run on the card; processing steps not natively supported by the T408 (such as composition and scaling) are handled by the CPU, with Norsk managing the handoff automatically. The T408 is suited to pure encode/decode workloads where on-card composition is not required.
Norsk detects available hardware at runtime and routes encode and decode tasks to it automatically. When a processing step is supported natively by the hardware, it runs on the card. When it is not, Norsk uses CPU for that step and returns the result to hardware for encode — without manual configuration of offload paths. The same workflow file runs on a CPU-only server, an NVIDIA GPU instance, or a NETINT Quadra machine.
Norsk has benchmarked three standard encoding ladders across CPU, NVIDIA T4, NVIDIA RTX 4000 Ada, and NETINT Quadra T1U hardware, for ultra-low latency, balanced, and high-quality encode profiles. CPU utilisation, GPU/VPU utilisation, and RAM usage are published for each combination — giving operators a concrete starting point for instance selection and cost modelling before deploying at scale.
Hardware acceleration reduces energy consumption per encoded channel relative to CPU-based encoding. Norsk is a founding member of the Greening of Streaming initiative, and the NETINT Quadra integration reflects a commitment to reducing the energy footprint of live streaming infrastructure. Higher channel density per server means fewer servers running, fewer watts consumed, and a smaller environmental footprint at scale.
Whether you are evaluating NETINT Quadra for an on-prem deployment, assessing GPU instances for cloud encoding, or want to understand the cost per channel difference at your specific scale, get in touch.
Talk to us