We just updated from 25.11.7 to 26.05.3 and notice an incorrect behavior for the state of nodes that are in a powered_down state. The state UNKNOWN is an error and ought to be IDLE for powered_down nodes. The sinfo command disagrees with itself about the node STATE: $ sinfo -N --long -n x008 Fri Aug 21 09:47:22 2026 NODELIST NODES PARTITION STATE CPUS S:C:T MEMORY TMP_DISK WEIGHT AVAIL_FE REASON x008 1 xeon24el8_test powered_dow 24 2:12:1 256000 140000 10424 xeon2650 none $ sinfo -t unknown -p xeon24el8_test PARTITION AVAIL TIMELIMIT NODES STATE NODELIST xeon24el8_test up 10:00 1 unk x008 The error is also reflected in "scontrol show node": $ scontrol show node x008 NodeName=x008 CoresPerSocket=12 CPUAlloc=0 CPUEfctv=24 CPUTot=24 CPULoad=0.00 AvailableFeatures=xeon2650v4,opa,xeon24,power_ipmi ActiveFeatures=xeon2650v4,opa,xeon24,power_ipmi Gres=(null) NodeAddr=x008 NodeHostName=x008 RealMemory=256000 AllocMem=0 FreeMem=N/A Sockets=2 Boards=1 State=UNKNOWN+POWERED_DOWN ThreadsPerCore=1 TmpDisk=140000 Weight=10424 Owner=N/A MCS_label=N/A Partitions=xeon24el8_test BootTime=None SlurmdStartTime=None LastBusyTime=Unknown ResumeAfterTime=None SuspendTime=3600 CfgTRES=cpu=24,mem=250G,billing=24 AllocTRES= CurrentWatts=0 AveWatts=0 This node was automatically powered off with the Slurm Power Saving feature as shown in the events table: $ sacctmgr show event Format=NodeName,TimeStart,Duration,State%-6,Reason%-40 where nodes=x008 NodeName TimeStart Duration State Reason --------------- ------------------- ------------- ------ ---------------------------------------- x008 2026-08-19T15:00:16 20:20:55 IDLE~ Powered down after SuspendTimeout x008 2026-08-20T11:21:11 00:14:49 UNK Powered down x008 2026-08-20T12:44:19 01:53:07 IDLE~ Powered down after SuspendTimeout x008 2026-08-20T17:33:28 08:17:21 IDLE~ Powered down after SuspendTimeout x008 2026-08-21T03:05:12 06:45:46 IDLE~ Powered down after SuspendTimeout There would seem to be a regression for nodes in the POWERED_DOWN state which should have the IDLE~ state. Thanks, Ole
Note added: We can't even select with sinfo those nodes that are in either idle+powered_down OR unknown+powered_down states as seen here: $ sinfo -t idle+powered_down,unknown+powered_down -N -n x008 NODELIST NODES PARTITION STATE We have to resort to only unknown+powered_down: $ sinfo -t unknown+powered_down -N -n x008 NODELIST NODES PARTITION STATE x008 1 xeon24el8_test unk When we wish to power up such "unknown" state nodes, we can resort to the new option "scontrol power up force <nodename>". I hope this behavior can be fixed soon. Thanks, Ole
I'd like to increase the priority to Medium because Slurm's power saving logic doesn't work correctly for partitions with "Unknown" node states. This blocks jobs on nodes which have a Powered_down state.
Hi Ole, Looking into this now. After running some routine tests on power saving in version 26.05.3, I still see nodes enter the "IDLE+POWERED_DOWN" state after they are run through the SuspendProgram. I'll need more information about your situation to understand what's going on. From a past ticket I see that you have "DebugFlags=Power" in your slurm.conf. If this is still true, please send along your slurmctld.log file. If you've unset this debug flag since then, please set it: > scontrol setdebugflags +power Let Slurm gather logs for an hour, then send along your slurmctld.log file. Afterwards you could unset the flag again: > scontrol setdebugflags -power Please also send along your up-to-date slurm.conf file and your SuspendProgram and ResumeProgram. Regards, Thomas
Hi Ole, I spoke too soon. I noticed that if I don't specify "State=" in my node configurations in slurm.conf, I see powered down nodes enter the "Unknown+powered_down" state after an "scontrol reconfigure": Before scontrol reconfig: > sinfo --Format=partition,nodes,nodelist,statecomplete:25,statelong --partition=defq > PARTITION NODES NODELIST STATECOMPLETE STATE > defq* 5 n[1-5] idle+powered_down idle~ After scontrol reconfig: > sinfo --Format=partition,nodes,nodelist,statecomplete:25,statelong --partition=defq > PARTITION NODES NODELIST STATECOMPLETE STATE > defq* 5 n[1-5] unknown+powered_down powered_down This occurs on 26.05+. This doesn't make sense, since, as you identified, the "Unknown" node state is just a placeholder state while the controller waits for the slurmd to respond. For a powered/powering down node, the slurmd won't respond, so an "Unknown+powered_down" node has to be manually started. Thank you for bringing this to our attention. I will look into a fix and keep you updated here. In the meantime, you could avoid this by specify a default state for your nodes in slurm.conf, for example "State=IDLE". Regards, Thomas
Hi Thomas, Thanks for looking into this regression and testing a workaround: (In reply to Thomas Sorkin from comment #6) > In the meantime, you could avoid this by specify a default state for your > nodes in slurm.conf, for example "State=IDLE". The slurm.conf manual doesn't list State=IDLE as an acceptable value: State of the node with respect to the initiation of user jobs. Acceptable values are CLOUD, DOWN, DRAIN, FAIL, FAILING, FUTURE and UNKNOWN. Node states of BUSY and IDLE should not be specified in the node configuration, but set the node state to UNKNOWN instead. Setting the node state to UNKNOWN will result in the node state being set to BUSY, IDLE or other appropriate state based upon recovered system state information. The default value is UNKNOWN. Nevertheless, I made the test of appending State=IDLE to the NodeName lines in slurm.conf and did "scontrol reconfig". Unfortunately, it seems that this KILLED a lot of jobs which were allocated to those nodes (I still need to confirm that these weren't killed for other reasons). So I quickly had to undo the State=IDLE configuration! Can you suggest other workarounds for this issue? Thanks, Ole
Note added: Can you change the behavior of State=Idle in slurm.conf so that it doesn't kill running jobs? This behavior was an unpleasant surprise for us. Thanks, Ole
Hi Ole, That is distressing and unexpected. Apologies for any inconvenience caused by this failed workaround. Do you mind sending along your slurmctld.log from this time? That can help me determine what exactly killed these jobs. I have a fix for avoiding the "unknown+powered_down" and "unknown+powering_down" node states in the works. I'm aiming to get it into our review process soon. I will keep you updated on its progress here. Regards, Thomas
Hi Thomas, (In reply to Thomas Sorkin from comment #9) > That is distressing and unexpected. Apologies for any inconvenience caused > by this failed workaround. Do you mind sending along your slurmctld.log from > this time? That can help me determine what exactly killed these jobs. I'll attach the slurmctld.log - can you please make the file private? I'm uncertain about what cancelled a number of jobs starting possibly about this log line: [2026-08-25T09:16:31.689] _job_complete: JobId=10770884 WEXITSTATUS 0 We reconfigured slurm.conf multiple times from 2026-08-25T09:16:23.817, so it's a little difficult to know when we had the State=IDLE and removed it again. As a temporary workaround, we disabled Slurm power saving by setting SuspendTime=Infinite for all partitions. > I have a fix for avoiding the "unknown+powered_down" and > "unknown+powering_down" node states in the works. I'm aiming to get it into > our review process soon. I will keep you updated on its progress here. Sounds good. We'd like a permanent fix for "idle+powered_down" nodes being incorrectly listed as "unknown". Let's see if this can be made in time for 26.05.4. Best regards, Ole
Thanks for sending the log, I just made it private. I will get back to you about when this fix will land. I just got it into our review process. Regards, Thomas
(In reply to Thomas Sorkin from comment #12) > Thanks for sending the log, I just made it private. > > I will get back to you about when this fix will land. I just got it into our > review process. Do you know if the fix was included in 26.05.4? Thanks, Ole
Hi Ole, Unfortunately, the changes could not go into 26.05.4. The latest Slurm releases of versions 26.05, 25.11, and 25.05 include fixes for several CVEs, so the amount of other bug fixes and functional changes we could introduce in them was quite limited. We expect many sites to upgrade to these dot releases to get the security patches, and we want to avoid the nightmare scenario of accidentally introducing an impactful bug on any of these security releases. See this public announcement for more details: > [1] https://groups.google.com/g/slurm-users/c/elYnXdOt3o0 I hope you can understand our reasoning here. On the upside, my fix for this issue is complete, reviewed, and ready to go. As soon as it is merged into 26.05.5, I'll point you to the necessary commit. Regards, Thomas
Hi Ole, The fix has now been merged. Here is the commit with the fix to avoid the "UNKNOWN+POWERED_DOWN" state: > [1] https://github.com/SchedMD/slurm/commit/6f83c4b Regards, Thomas
Hi Ole, I looked over your slurmctld log to examine the incident from comment #7. Upon further examination, I'm confident that what happened was not a bug and should not have had a huge impact, but I definitely gave you bad advice. Nothing in the logs suggests that running jobs were killed. They would have remained in StateSaveLocation, allowing your slurmctld to restore them. In fact, these two lines in the logs are where the slurmctld recovers job state and updates the state of allocated nodes: > [2026-08-25T09:16:24.027] Recovered information about 2753 jobs > ... > [2026-08-25T09:16:24.037] _sync_nodes_to_jobs updated state of 389 nodes Some pending jobs did fail on launch and were then requeued, since the slurmctld tried to allocate these on IDLE nodes that were actually suspended and powered down. I see 28 jobs being allocated in this window in lines 6707-6734 in the logs. These jobs were then requeued in lines 16531-16587 of the logs. With the fix merged in and the patch accessible in comment #15, I'll go ahead and close this ticket. Please don't hesitate to reopen it if you have any further questions. Regards, Thomas
Hi Thomas, Thanks very much for fixing this regression! Best regards, Ole