TrueNAS SCALE · post-production
Storage monitoring for TrueNAS post-production facilities
An editor reports a playback stall. Was a scrub running? Was a disk or pool degraded? Is the pool nearly full? Pillar keeps a shared timeline of what your TrueNAS SCALE system reported, so you can check storage first and decide what to do next.
Pillar shows evidence. It does not prevent stalls, prove what caused one, or change your pools for you.
The demo uses simulated data.
- Tested releaseTrueNAS SCALE 25.10.4, x86-64, in a virtual lab
- Not supportedTrueNAS CORE (FreeBSD)
- CollectionEvery 15 s by default: pools, scrubs, disks, capacity
What it covers
Three questions Pillar helps you answer
| Question | What Pillar shows | What you still check |
|---|---|---|
| Did storage maintenance overlap the stall? | Scrub and resilver state, progress and estimated completion from zpool status, next to your facility's operating hours. | Whether the maintenance actually caused the stall, and when to reschedule it. |
| Was a disk or pool struggling? | Pool and vdev health, error counters, SMART, and per-disk throughput, average wait and utilization for each collection interval. | The client, network, cache and application. Pillar does not see inside editing workstations. |
| Is capacity becoming a risk? | Pool and dataset capacity over time, with a runway estimate once a pool has at least 7 days of regular history. | Whether the estimate matches your project schedule. Estimates are uncertain and change with usage. |
Without the optional Observer, latency is the average wait per collection interval, not per-I/O percentiles. A short spike inside an interval can be averaged away. When data is unavailable, Pillar shows it as missing rather than guessing.
Example investigation
A playback-stall report during a scrub
This is a storage contention investigation prompted by a hypothetical playback report, reproduced on a disposable TrueNAS lab system with synthetic media. No editing application was tested. Each step is labeled as an observation, an inference or an action.
1. Symptom Observation
An editor reports that playback stalled at about 14:20. The operator opens that system's timeline for 14:00–14:40.
2. Evidence Observation
Pillar shows what the Agent recorded around that time: whether a scrub was scanning the pool, its progress, and the per-disk throughput and average wait in the same intervals.
3. Investigation Inference
If the scrub overlapped the report and disk wait rose with it, storage contention becomes a reasonable hypothesis. It is not proof: the network, the client and the application are still candidates, and a single report may be coincidence.
4. Decision Action
The operator gets approval to move the scrub schedule in TrueNAS outside edit hours, and asks editors to note the times of any further stalls. Pillar does not change the schedule.
5. Follow-up Observation
After the change, compare another observation window. Report an improvement only if the stall reports and the measurements actually change.
What the lab run measured
The run has a baseline window, a window with a scrub running, and a recovery window. Three recorders run the whole time and are kept separate: a fixed-rate reader standing in for one 24 fps stream (4 MiB per frame), TrueNAS's own zpool iostat, and the Pillar Agent.
Results not yet published
We'll publish the measured results here once the run has been reviewed, including a result that shows no contention. The limits below apply either way.
- A fixed-cadence reader is a storage-side stand-in for one playback stream. It is not an editing application and cannot show dropped frames in one.
- Virtual disks on a lab host do not behave like a production array of spinning disks or SSDs; timings will differ on real hardware.
- The scrub was started on purpose. The run shows whether the two overlapped and what each recorder saw, not that a scrub causes any particular facility's stalls.
Reproduction details
- TrueNAS SCALE 25.10.4, x86-64 virtual machine (2 vCPU, 8 GiB RAM) on a KVM lab host
- One RAIDZ1 pool of four 2 GiB virtual SATA disks, created through the TrueNAS middleware
- Dataset with compression off, 1 MiB records and primarycache=metadata, so reads reach the disks instead of ARC (a lab choice)
- Three 1 GiB clips of random (incompressible) synthetic data
- Pillar Agent running on the host, snapshot every 5 s; Observer not installed
- Windows: 90 s baseline, 120 s with a scrub running (restarted if it finishes early, stopped at the end of the window), 90 s recovery.
- A frame counts as late when its read finishes after its 1/24 s deadline.
- Every raw file is kept: per-frame timings,
zpool iostat -lat 1 s, each Agent snapshot, and scrub start/stop events.
Compatibility
What has been tested, and what hasn't
| Environment | Status | Notes |
|---|---|---|
| TrueNAS SCALE 25.10.4, x86-64 | Tested in a virtual lab | Pools (including boot-pool), members, datasets, zvols, encrypted datasets, capacity and degraded state matched what TrueNAS itself reported. Virtual disks, not physical hardware. |
| Other TrueNAS SCALE releases | Not yet verified | Pillar's installer detects TrueNAS SCALE, but detection is not certification: preflight reports it as “uncertain”. Tell us your release before you evaluate. |
| Physical HBAs, SES enclosures, NVMe | Not yet verified on TrueNAS | SMART and slot mapping depend on your controller and enclosure. Confirm them during evaluation. |
| TrueNAS CORE | Not supported | CORE runs FreeBSD. Pillar runs only on Linux. |
| arm64 | Not supported | x86-64 is the only certified architecture. |
Prerequisites
- An administrator account with shell access (System Settings → Shell, or SSH).
- The TrueNAS Apps service running, so Docker is available.
- Outbound HTTPS to Pillar (proxies are supported). No inbound ports.
- Approval from whoever owns the system to run a third-party container on it, and a maintenance window for the install.
Still being verified on TrueNAS
Staying running across reboots, updates and rollback, uninstall, and behavior across TrueNAS upgrades have not yet been demonstrated on TrueNAS. We will not describe them as supported until they have.
Agent and Observer
| Agent (required) | Observer (optional) | |
|---|---|---|
| Adds | Pool, scrub/resilver, disk, SMART, capacity and network telemetry | Block-latency and request-size histograms, including per-I/O percentiles |
| Needs | Docker, read access to block devices, SMART access | Linux 5.11 or newer with BTF, CAP_BPF and CAP_PERFMON, no network access |
| On TrueNAS | Collection tested on 25.10.4 (virtual lab) | Not yet tested on TrueNAS |
| If unavailable | No telemetry; the system shows as offline or unknown | The Agent keeps working; latency stays as per-interval averages |
Deployment and security
What installing Pillar involves
We haven't measured an install time on TrueNAS yet, so we don't quote one. Here is the work involved. Useful history, such as capacity runway, builds up over days after that.
- Confirm compatibility and approval. Check your release and hardware against the table above, and get sign-off to run a container on the system.
- Run and review the preflight. Download the read-only preflight script, read it, then run it. It reports the OS, architecture, block-device access and outbound connectivity.
- Generate a pairing command and install. A Facility or Organization Admin creates a one-time command in Settings → Agent pairing and runs it in the TrueNAS shell. The installer verifies the signed release and image before starting the Agent.
- Verify. Check that the system appears with a fresh heartbeat, and that the expected pools and drives and the available telemetry are all present.
- Configure facility context. Set the timezone, operating hours and alert routing before relying on notifications.
Access the Agent is given
What leaves the system
Outbound HTTPS only (TLS 1.3 for the Agent), with no listening ports. The Agent sends hostname, machine ID, OS, kernel and TrueNAS version, disk models, serials and SMART data, block and network counters, pool and dataset names and capacity, and a runtime posture report. It does not send file contents or file names. An optional setting partly redacts disk serials and MAC addresses; hostnames, pool names and dataset names are not anonymized. See Security and Data, retention & support.
Actions and removal
On TrueNAS, Pillar is observation only: it does not pause scrubs, blink LEDs or run commands on the system. Removal is pillar uninstall agent --yes --purge, followed by removing the system in the dashboard. Cancelling a subscription does not uninstall the Agent. Uninstall has not yet been run on TrueNAS, so plan to check it during evaluation.
Pricing and evaluation
What it costs and how to evaluate it
| Plan | Price | Appliances |
|---|---|---|
| Studio Fleet | $349/month, or $3,480/year | Up to 3 |
| Facility Enterprise | $799/month, or $7,920/year | Up to 10 |
Pricing is flat per plan, with no per-terabyte metering or per-seat charges. An appliance is one storage system running the Agent; the Observer on that same system is not billed separately. Studio Fleet includes 90 days of telemetry history. The trial lasts 14 days, covers 1 appliance and needs no card. It starts when the first system connects, not at sign-up, so waiting for approval doesn't use up trial days. To stop renewal, cancel in the billing portal at least 30 days before the end of your current subscription period. See Plans & billing.
An evaluation checklist
Decide what success looks like before you install. A good evaluation ends in a clear decision, even if the decision is that Pillar doesn't fit.
- The intended system connects, and the expected pools, drives and telemetry are present.
- A real or staged incident can be investigated from the shared timeline.
- Scrub or resilver overlap with your operating hours is visible, where your system runs them.
- Capacity information is useful once enough history has built up, and its uncertainty is clear.
- The installation, permissions and day-to-day running are acceptable to your facility.
See it with simulated data first
The demo shows the maintenance, storage and capacity views with simulated data. No account is needed.