# .NET Diagnostics Tooling — Server-Side Troubleshooting

How to install and use the official .NET diagnostic CLI tools (`dotnet-dump`, `dotnet-trace`,
`dotnet-counters`) directly on a GB5 host (e.g. the shared dev server, 217.216.78.142) when a
process is hung, slow, leaking memory, or burning CPU and the systemd/journalctl logs alone don't
explain why. Log-reading and bisection (commenting out code, redeploying, retrying) can rule things
out, but only a real dump or live trace shows the actual stuck thread/frame — reach for these first
once log-based investigation stalls, not as a last resort.

## Installation (no root/global install needed)

These tools install as regular .NET tools into any writable directory — no `sudo`, no touching the
system dotnet SDK install:

```bash
mkdir -p /tmp/.dotnet/tools
export DOTNET_ROOT=/root/.dotnet   # wherever this box's dotnet SDK actually lives
export PATH="$PATH:/tmp/.dotnet/tools"

dotnet tool install --tool-path /tmp/.dotnet/tools dotnet-dump
dotnet tool install --tool-path /tmp/.dotnet/tools dotnet-trace
dotnet tool install --tool-path /tmp/.dotnet/tools dotnet-counters
```

Persist the PATH addition for future sessions:

```bash
cat << 'EOF' >> ~/.bash_profile
# .NET Core diagnostic tools
export PATH="$PATH:/tmp/.dotnet/tools"
EOF
```

Requires outbound internet access from the box (NuGet.org) — confirmed available on
217.216.78.142 as of 2026-08-25.

## `dotnet-dump` — "why is this process stuck / crashed"

Best tool when a process is **hung** (not listening, not responding, no CPU) or you need to see
exactly which line of code a crash happened on days after the fact.

```bash
# Find the actual dotnet process (not the dapr sidecar wrapping it)
ps aux | grep FrameworkSL   # or whichever .dll

# Collect a full dump — needs sudo if the process runs as root (GB5's services do)
sudo /tmp/.dotnet/tools/dotnet-dump collect -p <PID> -o /tmp/hang.dmp

# Analyze non-interactively — the -c flag is required for scripted/piped use.
# Piping a command via `echo "pstacks" | dotnet-dump analyze ...` does NOT reliably work
# non-interactively (the REPL waits on stdin and the command hangs indefinitely under a
# backgrounded/piped shell) — always use `-c` instead:
sudo /tmp/.dotnet/tools/dotnet-dump analyze /tmp/hang.dmp -c pstacks -c exit
```

`pstacks` groups every thread by identical call stack — usually all you need. It shows both
managed and native frames, so a thread blocked inside `Microsoft.Data.SqlClient` or Kestrel's own
internals shows up by name, not just as a generic wait. Other useful `-c` commands:
`clrstack -all` (every managed thread's full stack, one at a time), `threadpool` (queue length —
the smoking gun for real ThreadPool starvation), `dumpheap -stat` (managed heap object counts, for
memory-growth investigations).

## `dotnet-trace` — CPU profiling / flame graphs while it's running

Best tool when a process is **alive but slow/CPU-hot**, and you want to see where time is actually
going without stopping it.

```bash
sudo /tmp/.dotnet/tools/dotnet-trace collect -p <PID> --duration 00:00:30 -o /tmp/trace.nettrace
```

Produces a `.nettrace` file — pull it to a local machine and open with `dotnet-trace convert` to
speedscope/chromium format, or Visual Studio/`PerfView` if available, for a flame graph.

## `dotnet-counters` — live CPU/memory/GC/ThreadPool metrics

Best tool for a quick, no-file, real-time read on a process's vitals — CPU %, working set, GC heap
size/gen0-2 counts, ThreadPool queue length and thread count, exception rate.

```bash
sudo /tmp/.dotnet/tools/dotnet-counters monitor -p <PID>
# or targeted:
sudo /tmp/.dotnet/tools/dotnet-counters monitor -p <PID> --counters System.Runtime,Microsoft.AspNetCore.Hosting
```

Ctrl+C to stop. Use this FIRST when the question is just "is this process actually under load or
truly idle/stuck" — it's the fastest of the three to reach for.

## Worked example — the `gb5.service` Kestrel-bind hang (2026-08-25)

Real incident this doc was written to capture, since it's the clearest illustration of why a dump
beats more log-reading:

**Symptom**: `gb5.service` (FrameworkSL) would start (background hosted services logging normally),
but Kestrel never bound to its configured port — Dapr logged "waiting for application to listen on
port 5000" indefinitely, `ss -tln` showed nothing bound, `strace -e trace=network` showed zero
`bind()` syscalls attempted. Intermittent — sometimes it came up fine, sometimes not, regardless of
which code was deployed (confirmed via a byte-for-byte-safe diff and an old known-good backup build
exhibiting the identical symptom).

**Ruled out via log-reading and bisection alone** (each took real effort, none found the cause):
MassTransit/RabbitMQ bus startup (disabled entirely — hang persisted), the entire feature-flagged
Scheduler/GOP/Qualifier code block (disabled — caused a different, unrelated DI crash instead),
disk space (was at 100% at one point but the symptom persisted even after it dropped to 92% free),
stale AppleDouble (`._*`) files in Dapr's components directory from a macOS-built tarball (real bug,
fixed, but not the cause of this hang), config drift in `AppPort`/`ASPNETCORE_URLS`.

**Found in under a minute with `dotnet-dump collect` + `pstacks`**: 4 threads blocked inside
`Microsoft.Data.SqlClient.SqlInternalConnectionTds.LoginNoFailover` → `Thread.Sleep` — SqlClient's
own connection-login retry-with-backoff loop. Kestrel's own `Heartbeat` timer thread was alive, but
no thread was executing `KestrelServerImpl.BindAsync` at the moment of the dump — it was simply
waiting its turn in ASP.NET Core's sequential `Host.StartAsync()` hosted-service pipeline, behind
other hosted services stuck retrying slow SQL logins. Cross-checked against SQL Server itself (172
live sessions, 116 connections, far under the 32767 max) — real transient login/negotiation
slowness under concurrent load from many hosts sharing one dev SQL Server instance, not a connection
-count ceiling.

**Architectural implication** (tracked as a real follow-up, not yet fixed — see
[[project_metadata_promotion_gop_hardening]] and the plan doc): `FrameworkSL`/`gb5.service` has
noticeably more hosted services opening their own DB connections at boot than the other GB5 Hosts
do (Scheduler, GOP worker, multiple Orphan-cleanup jobs, WorkflowAutoApprove, DataSync, Action
processor, etc.), and ASP.NET Core's Generic Host starts registered `IHostedService`s sequentially —
so any one of them being slow to open its first connection delays every hosted service registered
after it, including Kestrel's own binding. Under real production load (not just a shared dev box)
this same mechanism would make `gb5.service` itself unavailable for however long the slowest
hosted-service's DB login takes, which could be substantial in a high-concurrency environment.
Correcting this — decoupling Kestrel's own bind/listen from custom hosted-service startup order, and
auditing every `IHostedService.StartAsync` override in `FrameworkSL`/`FrameworkBLL` for anything that
blocks on I/O rather than just kicking off a background `Task` (the standard, non-blocking
`BackgroundService` pattern) — is real, important work for high-volume production readiness, but
deliberately scoped as a separate follow-up session rather than a quick patch made mid-investigation.
