Skip to main content
SCAIRE

Diagnostic Report #74

Specialist assessments and Faire's synthesized analysis.

scheduled 2026-07-23 7:04:04 AM run_id: 1f196358-b82c-4eec-b047-e88d8f009169

Faire ยท Synthesis

๐ŸŸข Bottom line

All green, Boss. Three domains, three healthy verdicts, zero service failures, zero disk errors, zero kernel issues. 71 LXC containers and 3 VMs humming along on 56 cores with loads under 2. Nothing needs your attention right now โ€” there are a few items worth a look during your next maintenance window, but nothing urgent.


๐ŸŸข Compute (Grian)

type: status
title: Compute health
items: [{lb: CPU, st: ok}, {lb: Memory, st: ok}, {lb: Services, st: ok}, {lb: Kernel, st: ok}, {lb: Instances, st: ok}]
Metric Value Status
Load avg (1/5/15) 1.05 / 1.31 / 1.65 on 56 cores โœ… Idle
CPU usage 6.05% (2.10% usr, 2.35% sys) โœ…
Memory 555,789 MB / 773,913 MB (71.8%) โœ… Comfortable
PSI CPU / Mem / IO 0.03 / 0.00 / 0.36 โœ… No contention
Failed systemd units 0 โœ…
Kernel errors 0 โœ…
GPU None detected N/A
Instances 71 LXC + 3 VMs, all stable โœ…

Top resource consumers โ€” all expected services, no runaways:

Process CPU Memory Notes
VM 144 (dockercompose1) 16.1% 0.8% KVM VM, Docker workloads
sonarr 10.6% 0.0% Media organizer
jellyfin 5.4% 0.1% Media streaming
qdrant 1.0% 0.3% Vector database
  • โœ… PSI memory at 0.00 โ€” absolutely zero memory pressure despite 71.8% utilization.
  • โš ๏ธ PSI IO at 0.36 โ€” the highest of the three PSI values, but still well into "no problem" territory. Worth a passing glance if it trends upward.
  • โš ๏ธ Memory at 71.8% โ€” Grian recommends watching for the ~80% threshold. Plenty of headroom today.

๐ŸŸข Storage (Saor)

type: status
title: Storage health
items: [{lb: data pool, st: ok}, {lb: rpool, st: ok}, {lb: rustpool, st: ok}, {lb: SMART, st: ok}, {lb: Scrub, st: ok}, {lb: ARC, st: ok}]

ZFS Pools:

Pool Size Used Free Cap % Health Frag %
data 18.2T 11.2T 6.97T 61% ONLINE 39%
rpool 888G 226G 662G 25% ONLINE 49%
rustpool 2.45T 317G 2.14T 12% ONLINE 16%

ARC Cache:

Metric Value Status
Hit ratio 99.24% (threshold >80%) โœ… Excellent
Cache size ~500 GB / 512 GB max โœ… Near-full, working as intended

I/O Latency:

Pool / Media Read await Write await Threshold Status
data (HDD) 0.57โ€“0.97 ms 0.60โ€“0.76 ms <10 ms โœ…
rpool (SSD) 6.70โ€“6.76 ms 9.30โ€“9.46 ms <5โ€“10 ms โš ๏ธ Borderline
  • โœ… All disks ONLINE, 0 read/write/checksum errors across all pools.
  • โœ… Scrub completed Jul 12, 2026 โ€” 0B repaired, 0 errors.
  • โœ… All 5 Proxmox storage backends ACTIVE (local, localZFS, localbackup, rustbackup, rustpool).
  • โš ๏ธ rpool write await at 9.30โ€“9.46 ms โ€” technically within the <10 ms threshold but grazing the ceiling. Combined with the 49% fragmentation on this boot-pool mirror, it's a heads-up rather than a fire. A scrub or dump/restore during planned maintenance would clean it up.
  • โš ๏ธ data pool at 61% โ€” 6.97 TB free. Set an alert at 80% to plan expansion.

๐ŸŸข Network (Nasca)

type: status
title: Network health
items: [{lb: Links, st: ok}, {lb: DNS, st: ok}, {lb: Gateway, st: ok}, {lb: Errors, st: ok}]
Metric Value Status
Interfaces UP nic0, nic2, bond0, vmbr0 โœ…
IP (vmbr0) 192.168.1.34/24 โœ…
Gateway 192.168.1.1 via vmbr0 โ€” reachable โœ…
DNS 7 lookups resolved to 192.168.1.2 โœ…
Link speeds nic0/nic2: 10 Gb/s, bond0/vmbr0: 20 Gb/s (full duplex) โœ…
Utilization All interfaces <0.5% โœ…
vmbr0 rx errors 53.2 err/min (0.89 err/sec) vs 100 err/sec threshold โœ… Acceptable
Firewall 0 rules, 0 chains โš ๏ธ See note
  • โœ… All links up, low utilization, DNS functional, gateway reachable.
  • โš ๏ธ vmbr0 rx errors at 0.89 err/sec โ€” well below the 100 err/sec threshold and almost certainly benign (could be port flooding, broadcast noise, etc.), but it's non-zero. Worth watching for trend.
  • โš ๏ธ No firewall rules active โ€” 0 rules, 0 chains. May be intentional (upstream firewall, dedicated security appliance), but if this node is internet-facing or on a flat network, that's a gap worth confirming.

Recommended actions

All low-priority. Nothing urgent.

# Action Priority Needs confirmation?
1 Schedule rpool scrub or dump/restore to address 49% fragmentation Low โœ… Safe โ€” zpool scrub rpool
2 Set capacity alert for data pool at 80% (currently 61%, 6.97 TB free) Low โœ… Safe โ€” configure in your monitoring stack
3 Confirm whether 0 firewall rules on this node is intentional Low โš ๏ธ Yes โ€” your call, Boss
4 Keep an eye on memory trending toward 80% and PSI IO if it climbs past 1.0 Low โœ… Automatic monitoring

That's it. All green, nothing on fire. The kind of report that makes for a quiet day ๐Ÿ˜Œ.

Specialist reports (3)

Grian 152.1s

I'll begin the compute & system resources diagnostic run by executing the diagnostic script.

Let me gather some additional context on the processes to ensure nothing is being missed.

Compute & System Resources Assessment

Overall Health Status: HEALTHY

Key Metrics

Metric Value Status
CPU Load (1/5/15 min) 1.05 / 1.31 / 1.65 โœ… Excellent (56 cores available)
CPU Usage 6.05% (2.10% usr, 2.35% sys) โœ… Minimal utilization
Memory Usage 555,789 MB / 773,913 MB (71.8%) โœ… Normal range, no pressure
PSI CPU 0.03 โœ… No resource contention
PSI Memory 0.00 โœ… Zero memory pressure
PSI IO 0.36 โœ… Minimal IO pressure
Failed Systemd Units 0 โœ… All services healthy
Kernel Errors 0 โœ… No kernel issues
GPU None detected N/A
Proxmox Instances 71 LXC + 3 VMs โœ… Stable

Top Resource Consumers

  • VM 144 (dockercompose1): 16.1% CPU, 0.8% memory - KVM VM with Docker workloads
  • qdrant: 1.0% CPU, 0.3% memory - Vector database service
  • jellyfin: 5.4% CPU, 0.1% memory - Media streaming service
  • sonarr: 10.6% CPU, 0.0% memory - Media organizer

All processes are expected services running within LXC containers/VMs with normal resource consumption.

Anomalies & Concerning Findings

None detected. All compute and system resource metrics are within healthy ranges. The system has substantial headroom (56 cores, 28% free memory) and shows no signs of resource contention, service failures, or kernel errors.

Recommended Actions

No immediate actions required. Continue monitoring memory usage as it approaches ~80% threshold, but current utilization is well within safe operating parameters.

[SEVERITY: healthy]

Saor 164s

I'll begin the storage and filesystem diagnostic run. Let me start by running the diagnostic script.

Let me gather additional context on ARC cache statistics and disk health to complete the assessment.

Storage & Filesystem Assessment

Overall Health Status: HEALTHY


Key Metrics

ZFS Pools:

Pool Size Used Free Capacity Health Fragmentation
data 18.2T 11.2T 6.97T 61% ONLINE 39%
rpool 888G 226G 662G 25% ONLINE 49%
rustpool 2.45T 317G 2.14T 12% ONLINE 16%

ARC Cache:

  • Hit ratio: 99.24% (excellent; threshold >80%)
  • Cache size: ~500GB / 512GB max

I/O Performance:

  • HDD await (data pool): 0.57-0.97ms read, 0.60-0.76ms write (threshold: <10ms)
  • SSD await (rpool): 6.70-6.76ms read, 9.30-9.46ms write (threshold: <5-10ms)
  • Disk utilization: All under 12% (threshold: >70%)

Disk Health:

  • All disks: ONLINE
  • Read/Write/Checksum errors: 0
  • Scrub status: Completed 0B repaired, 0 errors (Jul 12, 2026)

Proxmox Storage Backends:

  • All 5 backends (local, localZFS, localbackup, rustbackup, rustpool): ACTIVE

Findings

No Critical Issues:

  • All ZFS pools are ONLINE with no errors
  • No disk failures or SMART alerts
  • Scrub completed successfully across all pools
  • Excellent ARC cache performance
  • I/O latency well within acceptable limits
  • Capacity utilization healthy (data at 61% should be monitored)

One Observation:

  • rpool fragmentation at 49% is elevated. This is not causing immediate performance issues (I/O remains good), but as this is the boot pool on mirrored SSDs, sustained high fragmentation could impact future performance. A scrub or periodic optimization would be beneficial when maintenance allows.

Recommendations for Faire

  1. Monitor data pool capacity - currently at 61% (6.97TB free). Set alert threshold at 80% to allow time for expansion.
  2. Schedule rpool optimization - the 49% fragmentation on the boot pool (mirror) could be improved via a scrub and potentially a temporary dump/restore cycle during planned maintenance.
  3. Continue regular scrubs - the last successful scrubs were ~11 days ago; maintain monthly scrub cadence.

[SEVERITY: healthy]

Nasca 131.6s

I'll begin the network infrastructure diagnostic run. Let me execute the diagnostic script first.

Network Infrastructure Assessment

Overall Health Status: HEALTHY

Key Metrics

Metric Value Status
Interface Links nic0, nic2, bond0, vmbr0 all UP โœ“ Normal
IP Address vmbr0: 192.168.1.34/24 โœ“ Configured
Gateway 192.168.1.1 via vmbr0 โœ“ Reachable
DNS Resolution 7 lookups resolved to 192.168.1.2 โœ“ Working
Interface Utilization All interfaces <0.5% (LOW) โœ“ Normal
Interface Errors vmbr0: 53.2 rx err/min (0.89 err/sec) โœ“ Acceptable (<100 err/sec threshold)
Link Speeds nic0/nic2: 10Gb/s, bond0/vmbr0: 20Gb/s โœ“ Full duplex
Firewall 0 rules, 0 chains No active filtering

Anomalies / Concerning Findings

None detected. The diagnostic shows:

  • All network interfaces are UP with proper link status
  • Utilization is low across all interfaces (all under 0.5%)
  • Error rates are well below critical thresholds (0.89 errors/sec vs 100 err/sec threshold)
  • DNS resolution is functional
  • Gateway is configured and reachable

Observation: Firewall shows no active rules (0 rules, 0 chains). This may be intentional depending on your security model, but in production environments, consider whether firewall rules should be in place for network segmentation and security.

Recommended Actions for Faire

No immediate action required. Continue monitoring network utilization trends. Consider reviewing firewall policy requirements if network security hardening is desired.

[SEVERITY: healthy]

Back to chat