KNODES(1)toolbelt manualKNODES(1)

knodes

Show node health and pressure at a glance.

Synopsis

knodes [NODE] [-- kubectl options]

Description

kubectl get nodes says Ready or NotReady and little else. A node can be Ready and still be in trouble. Under MemoryPressure, DiskPressure or PIDPressure the kubelet refuses new pods and evicts running ones, and get nodes shows none of that. You find it in kubectl describe node, one node at a time, in a long block of conditions.

knodes prints one row per node with its status, roles, age and kubelet version. A CONDITIONS column appears when any node has a problem, and says ok for the others. When every node is fine, the table ends with all conditions normal.

A taint on a node keeps pods off it unless they tolerate it, which is a common reason a node gets no pods. Each taint gets a line under the table, such as taint k8s-worker-1 dedicated=db:NoSchedule. Two kinds are left out. The control-plane taint on a node with the control-plane role is expected there. The node.kubernetes.io/ taints that Kubernetes adds for NotReady, pressure or a cordon repeat what the table already shows.

Each problem then gets a short block under the table.

Name a node to see only that one.

knodes only reads. It never cordons, drains or changes a node.

Options

OptionWhat it does
-q, --quietPrint only the table, no taints, no detail blocks and no closing line.
-v, --verbosePrint each kubectl command before it runs.
-h, --helpShow the help.

What counts as a problem

ConditionA problem whenWhat it means
Readynot TrueThe kubelet is down, or has not reported in time. Shown as NotReady.
MemoryPressureTrueThe node is low on memory. The kubelet evicts pods, those most over their memory request first.
DiskPressureTrueThe node is low on disk or inodes. The kubelet deletes unused images and then evicts pods.
PIDPressureTrueThe node is running out of process ids.
NetworkUnavailableTrueThe network plugin has not set up the node's network.

The STATUS column also shows SchedulingDisabled for a cordoned node, as kubectl get nodes does. The ROLE column lists the node-role.kubernetes.io/ labels, or worker for a node with none.

Pass-through

Options after -- go to every kubectl call.

knodes -- --context kind-kind

Needs

kubectl and jq. metrics-server, optional, for the top line. Without it, the line is left out and the rest still prints.

The account behind your context needs list on nodes, and on events and pods in every namespace for the detail lines.

Examples

A node under memory pressure

knodes
NODE           STATUS  ROLE           AGE  VERSION  CONDITIONS
k8s-master     Ready   control-plane  41d  v1.35.2  ok
k8s-worker-1   Ready   worker         41d  v1.35.2  ok
k8s-worker-2   Ready   worker         41d  v1.35.2  MemoryPressure
k8s-worker-2   MemoryPressure since 6h
  kubelet  has insufficient memory available
  evicted  2 pods in the last hour
  top      prometheus-0 1.9Gi, loki-0 820Mi, api-2wq8d 210Mi
echo $?
1

prometheus-0 is the pod to look at. kres --pods k8s-worker-2 shows it uses far more than it requests.

A node that stopped reporting

knodes
NODE           STATUS                       ROLE           AGE  VERSION  CONDITIONS
k8s-master     Ready                        control-plane  41d  v1.35.2  ok
k8s-worker-1   NotReady,SchedulingDisabled  worker         41d  v1.35.2  NotReady
k8s-worker-2   Ready                        worker         41d  v1.35.2  ok
k8s-worker-1   NotReady since 25m
  message  Kubelet stopped posting node status.

The node is down or cut off from the API server. After 5 minutes by default, its pods are marked for eviction and their owners start copies on the other nodes.

A healthy node

knodes k8s-master
NODE         STATUS  ROLE           AGE  VERSION
k8s-master   Ready   control-plane  41d  v1.35.2
all conditions normal
echo $?
0

A tainted node

knodes
NODE                     STATUS  ROLE           AGE  VERSION
toolbelt-control-plane   Ready   control-plane  40m  v1.37.0
taint  toolbelt-control-plane  dedicated=db:NoSchedule
taint  toolbelt-control-plane  gpu:PreferNoSchedule
all conditions normal

A taint is not a problem, so the exit status stays 0. Pods that should run there need a matching toleration, and kubectl taint node NAME dedicated- removes the taint.

Only the table

knodes -q
NODE           STATUS  ROLE           AGE  VERSION  CONDITIONS
k8s-master     Ready   control-plane  41d  v1.35.2  ok
k8s-worker-1   Ready   worker         41d  v1.35.2  ok
k8s-worker-2   Ready   worker         41d  v1.35.2  MemoryPressure

What it runs

knodes -v > /dev/null
+ kubectl get nodes -o json
+ kubectl get events -A --field-selector reason=Evicted -o json
+ kubectl top pods -A --no-headers
+ kubectl get pods -A --field-selector spec.nodeName=k8s-worker-2 -o json

The last three run only when a node is under pressure.

A typo in the node name

knodes nope
knodes: no node nope. Nodes are k8s-master, k8s-worker-1 and k8s-worker-2

Troubleshooting

evicted says 0 pods, but pods were evicted
Events expire after an hour by default, and some clusters keep them for less. An eviction older than that leaves only the pod itself with the status Evicted, which kclean lists.
No top line
metrics-server is not installed or not ready. kres says which.
A kind node is NotReady right after kind create cluster
The network plugin takes a few seconds to start. Wait and run knodes again.

Exit status

CodeMeaning
0Every node looked at is Ready with no pressure.
1A node is NotReady or has a problem, no such node, or kubectl failed.
2Bad usage, such as two node names.
3kubectl or jq is missing.

See also

kres, kclean, kwhy, mem, kubectl-describe(1)