P
PILLARDOCS
Metric Manual

Incident Attribution & Metric Interpretation

When an edit bay stutters during playback, Technical Directors must quickly differentiate between client workstation strain, network saturation, and backend storage tail latency. This manual covers the core metrics monitored by Pillar.

1. Drop-Frame Root-Cause Attribution & Latency Budgets

Use the interactive tool below to correlate storage read await latencies against video delivery deadlines:

Interactive Playback Budget Tool

Video Frame Delivery Window vs. Disk Await

Formula: Budget (ms) = 1000 / FPS

Format Profile: Narrative Feature / Netflix Master

Simulated Single-Drive p99 Read Await:25 ms
1ms (NVMe Tier)41.7ms (24fps Budget)120ms (Failing Spindle ERC)
Playback Frame Delivery Budget41.7 ms

Storage must deliver each frame block within this window to maintain seamless NLE playback buffer.

Editorial Impact Assessment✓ SEAMLESS BUFFER CONTINUITY

Read await (25ms) is safely within the 41.7ms deadline (60% consumed).

Standard Video Playback Delivery Windows

Timeline FormatTarget FPSFrame Delivery WindowPillar Alert ThresholdSeverity
Narrative Feature / Netflix Master23.976 / 24.0 fps41.7 msp99 await > 40.0 msCRITICAL
Broadcast Television / Commercial29.97 / 30.0 fps33.3 msp99 await > 30.0 msCRITICAL
Sports Broadcast / High Frame Rate59.94 / 60.0 fps16.6 msp99 await > 15.0 msCRITICAL
Virtual Production / In-Camera VFX119.88 / 120.0 fps8.3 msp99 await > 7.5 msFATAL

Verify real-time host await latencies directly from the terminal:

bash
iostat -xz 1 5 | awk '{print $1, $4, $10, $14}'

2. ZFS Copy-on-Write (CoW) 80% Allocation Cliff

ZFS pools experience a dramatic performance drop when utilization exceeds 80%. This is caused by a change in how the ZFS kernel module searches for free disk space:

Below 80% Capacity: First-Fit

The kernel grabs the first contiguous block in the metaslab space map. Allocation requires minimal CPU cycles and write latencies remain low (1–5 ms).

Above 80% Capacity: Best-Fit

Space maps become fragmented. ZFS switches to best-fit allocation, traversing in-memory and on-disk AVL trees. Latency spikes up to 150–300 ms, causing playback dropped frames.

Check your current pool fragmentation index and capacity:

bash
zpool list -o name,size,alloc,free,cap,frag

3. Scrub Windowing & Resilver Guardrails

A ZFS scrub validates block checksums to detect silent bit rot. However, on large multi-terabyte arrays, an unconstrained scrub competes directly with editorial playback:

  • 25% to 45% Throughput Degradation: Concurrent 4K ProRes playback streams drop as scrub I/O occupies disk queues.
  • Window Overrun Tracking: Pillar forecasts the estimated scrub completion time (eta_seconds). If a weekend scrub is projected to continue past Monday 08:00, Pillar fires a preemptive advisory alert.
  • Dynamic Priority Throttling: Engineers can reduce scrub impact by adjusting zfs_scrub_delay or pausing the pass during critical client review sessions.