System and Infrastructure Status News

Anvil outage

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: anvil.purdue.access-ci.org, anvil-gpu.purdue.access-ci.org

Start Date: October 13, 2024, 1:00 p.m.

End Date: October 13, 2024, 9:00 p.m.

Update: As of 5:00PM EST Oct 13, 2024, most compute nodes have been brought back online and Anvil has been back to normal operations. If you had jobs running around 09:00AM when the power interruption occurred, those jobs were terminated, so you may need to check the state of running jobs that were dependent upon work that was interrupted. Original Post: Shortly before 9:00AM EST Oct 13, 2024, Anvil has experienced a power interruption due to a campus power outage. It has caused compute nodes down and scheduling paused. Our engineers have arrived on campus and found additional impacts from today's power interruption. At of 12:18PM EST, scheduling has been resumed and we are still working on bringing remaining nodes back. We are at reduced capacity now but users should see stable systems at the moment. Additional efforts to restore full service are undertaken. Will post updates before the end of today. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket if you have any questions.

Posted: March 20, 2026 • Author: Guangzhen Jin

ACES Partial Maintenance, October 9-10

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: aces.tamu.access-ci.org

Start Date: October 9, 2024, 2:00 p.m.

End Date: October 10, 2024, 11:00 p.m.

There will be upcoming maintenance on some GPU components for the ACES cluster. The maintenance period is 9am on Wednesday October 9 to 6pm on Thursday October 10. Other parts of ACES will remain available. For specific details, please visit the HPRC website at https://hprc.tamu.edu/.

Posted: March 20, 2026 • Author: Francis Dang

ACCESS XDMoD Downtime

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: xdmod.access-ci.org

Start Date: September 30, 2024, 12:00 p.m.

End Date: October 3, 2024, 5:00 p.m.

Update 10/01/2024: The main web service has successfully been updated, however there was an unrelated power outage that is causing a hardware issue for one of the databases. The following services are only partially available at this time: - Efficiency tab - Job Viewer tab for the SUPReMM realm Another update will be posted when the service is fully restored in the coming days. We apologize for any inconveniences. End of Update There is a scheduled downtime for ACCESS XDMoD on Monday, September 30th from approximately 07:00 EDT until 07:00 EDT the following day. The service will be updated from XDMoD 10.5 to 11.0 during this time. The web service and API access through the Data Analytics Framework will be unavailable during the outage. If you wish to use the xdmod-data Python package with ACCESS XDMoD after the upgrade, you will need to upgrade the package to the latest 1.0.1 version using pip install --upgrade xdmod-data. An update will be posted when the service is fully restored. We apologize for any inconvenience this may cause.

Posted: March 20, 2026 • Author: Conner Saeli

DUO Authentication Maintenance Saturday September 7 2024 05:00am EDT

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: allocations.access-ci.org, identity.access-ci.org, registry.access-ci.org, access-ci.org

Start Date: September 7, 2024, 9:00 a.m.

End Date: September 7, 2024, 3:00 p.m.

DUO Security has announced a maintenance period for Saturday September 7 starting at 05:00am EDT, for an estimated period of 6 hours, which may cause disruption to ACCESS DUO authentication attempts: Begin forwarded message: From: support-noreply@status.duosecurity.com Subject: Duo Maintenance - Multiple Deployments: Scheduled Maintenance - 7 September 2024 Date: August 30, 2024 at 4:45:20 PM EDT Multiple Deployments: Scheduled Maintenance Upcoming scheduled maintenance noticeThe Duo Site Reliability Engineering (SRE) team is scheduling regular maintenance on the following deployments, beginning with DUO70: DUO70 DUO1 DUO55 DUO62 DUO63 DUO73 DUO65 DUO79 DUO77 DUO80 DUO78 The goal of this maintenance is to better balance authentication traffic across all Duo deployments in our continued efforts to provide a high performance and resilient service. We expect minimum user impact since each specific migration window will occur during non-peak times for your organization. Do I need to take action? If your organization has strict IP filtering/firewall rules in place, you should ensure that all Duo IP ranges listed here https://help.duo.com/s/article/1337 and DNS names exist in any filtering rules if firewalls, SSL inspection, or other proxy rules whitelist communication to specific Duo IP ranges. In the event of an outage or failure, Duo’s service could automatically failover to any IP in those ranges. Otherwise, no action is required. How does this migration affect you? During the migration, we expect minimum intermittent authentication failures when users attempt to log in to Duo protected applications. Users may need to retry the authentication or potentially wait a few minutes before attempting another login. After the migration Duo Administrators with the Owner and Administrator role will have their notification settings for the Duo Status Page https://status.duo.com/ adjusted to reflect the new deployment within 24 hours. Duo Administrators with the other roles will need to update their Duo Status Page preferences manually using the steps in https://help.duo.com/s/article/2060. Resources • Documentation: How do I find my StatusPage deployment ID in the Duo Admin Panel and sign up for updates? - https://help.duo.com/s/article/2060 * Article: What are Duo's IP ranges and data residency areas by deployment? - https://help.duo.com/s/article/1337 * Duo Status Page - https://status.duo.com/ For any questions or concerns please email support@duo.com (mailto:support@duo.com). Thank you for being a Cisco Duo customer! Duo Site Reliability Engineering Start time Sep 7, 05:00 EDT Estimated duration 6 hours Components affected DUO1 - Core Authentication Service DUO1 - Admin Panel DUO1 - Push Delivery DUO1 - Phone Call Delivery DUO1 - SMS Message Delivery ...and 72 more components. View full scheduled maintenance details (https://stspg.io/dx573gm8pvdm)

Posted: March 20, 2026 • Author: Derek Simmel

Expanse maintenance - 5AM (PT) 07/24/2024 to 5AM (PT) 07/25/2024

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org

Start Date: July 24, 2024, 12:00 p.m.

End Date: July 25, 2024, 12:00 p.m.

SDSC Expanse will be under maintenance from 5AM (PT) 07/24/2024 to 5AM (PT) 07/25/2024. During this time we will be working on the direct liquid cooling (DLC) infrastructure. A reservation has been put in place to prevent jobs from running during the maintenance period. The "squeue" command output will show "ReqNodeNotAvail, Reserved for maintenance" for jobs that do not fit in the time period before the maintenance begins. These jobs will run after we release the maintenance reservation.

Posted: March 20, 2026 • Author: Mahidhar Tatineni

Kerberos Replica Potential Outage

Published

Infrastructure News Type: Degraded

Affected Infrastructure: kerberos.access-ci.org

Start Date: July 18, 2024, 9:12 p.m.

End Date: July 20, 2024, 10:00 a.m.

Due to a cooling issue at NCSA, it is possible that the replica KDC hosted on-site could become unresponsive. Kerberos services should still operate in the event the replica goes down but their may be a delay as things fail over to another replica server.

Posted: March 20, 2026 • Author: Jacob Gallion

CILogon logins are failing for some users

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: identity.access-ci.org

Start Date: July 15, 2024, 3:07 p.m.

End Date: July 15, 2024, 5:28 p.m.

Web login to ACCESS portals and other services is failing for some users due to a problem in the CILogon service. Technical details: The AWS EFS mount points used by the LDAP servers became unavailable. We are in the process of removing/replacing EFS mount points and restarting the LDAP servers. The process will take several minutes.

Posted: March 20, 2026 • Author: John-Paul Navarro

Jira Service Management Incident - Some products are hard down - 3 July 2024

Published

Infrastructure News Type: Degraded

Affected Infrastructure: tickets.access-ci.org

Start Date: July 3, 2024, 11:42 p.m.

End Date: July 4, 2024, 1:43 a.m.

Between 03-07-2024 20:08 UTC to 03-07-2024 20:31 UTC, we experienced issue creation and project for Jira Service Managements. The issue has been resolved and the service is operating normally. https://jira-service-management.status.atlassian.com/incidents/ctpbgx33bd7r

Posted: March 20, 2026 • Author: Dinuka De Silva

Important Update: Changes to Ticket Automation and Status Updates

Published

Infrastructure News Type: Reconfiguration

Affected Infrastructure: tickets.access-ci.org

Start Date: June 29, 2024, 6:19 p.m.

End Date: July 1, 2024, 5:00 a.m.

Due to recent licensing changes implemented by Atlassian, we have had to re-configure how automated ticket actions are performed. As a result, the status changes based on the comments added to tickets won’t be happening anymore. Additionally, some minor automatic corrections will no longer happen instantaneously but will happen eventually. These changes are being implemented over the next few days, so please bear with us if you encounter any discrepancies. We expect these changes to be completed by July 1, 2024.

Posted: March 20, 2026 • Author: Dinuka De Silva

Ticketing System Automation Rules Currently Not Functioning

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: tickets.access-ci.org

Start Date: June 21, 2024, 4:00 a.m.

End Date: June 21, 2024, 6:01 a.m.

Hi Everyone, We have reached the limit of automation rule executions on our ticketing system. As a result, you may miss notifications, and auto-generated fields will not be populated. However, ticket creation and management (queues, etc.) through the portal are functioning as usual. Please bear with us as we work to resolve this issue as soon as possible. Thank you for your understanding. Regards, Dinuka (On behalf of ACCESS Operations)

Posted: March 20, 2026 • Author: Dinuka De Silva

ACES Partial Unavailability, May 29-30

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: aces.tamu.access-ci.org

Start Date: May 29, 2024, 2:00 p.m.

End Date: May 31, 2024, 1:00 a.m.

ACES will be partially available during May 29-30 as we repair additional hardware issues for the PCIe composability fabrics.

Posted: March 20, 2026 • Author: Francis Dang

ACCESS XDMoD Downtime

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: xdmod.access-ci.org

Start Date: May 23, 2024, 1:00 p.m.

End Date: May 24, 2024, 1:00 p.m.

Update 05/24/24: ACCESS XDMoD is fully functional now. All services are available at this time. There will be a full outage for ACCESS XDMoD today from approximately 09:00 EDT on 5/23 until 09:00 EDT on 5/24. The web service and API access through the Data Analytics Framework may be unavailable at points during this time. An update will be posted when the service is fully restored.

Posted: March 20, 2026 • Author: Conner Saeli

SDSC Expanse Maintenance 8AM-8PM (PT), Monday, May 20, 2024

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org

Start Date: May 20, 2024, 3:00 p.m.

End Date: May 21, 2024, 3:00 a.m.

Dear Expanse User, We will have a maintenance period on Expanse 8AM-8PM (PT), May 20, 2024. During this maintenance, we will be performing kernel updates and also work on the Ceph filesystem. We have a reservation in place to prevent jobs from running during this period. The "squeue" output will show "ReqNodeNotAvail, Reserved for maintenance" for jobs that do not fit in the time period before the maintenance begins. These jobs will run after we release the maintenance reservation. Thanks SDSC User Support Staff

Posted: March 20, 2026 • Author: Mahidhar Tatineni

Anvil Cluster Maintenance - Partial

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: anvil.purdue.access-ci.org

Start Date: May 13, 2024, 12:00 p.m.

End Date: May 14, 2024, 12:00 p.m.

The Anvil system will have reduced capacity Monday, May 13 from 8:00am - Tuesday, May 14 at 8:00am EDT for scheduled maintenance. Some compute nodes will be powered off during the maintenance for some power work. How does this maintenance impact you? - This is a standard partial maintenance for Anvil system - OS and Slurm will still be functioning with reduced capacity - User jobs might experience longer waiting time - No user impact is expected once Anvil is returned to service Anvil will return to full production by Tuesday, May 14, 2024 at 8:00am EDT. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/open-a-ticket if you have any questions.

Posted: March 20, 2026 • Author: Guangzhen Jin

ACCESS XDMoD Downtime

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: xdmod.access-ci.org

Start Date: April 29, 2024, 3:00 p.m.

End Date: May 1, 2024, 10:00 p.m.

UPDATE: 5/2/24 The service has been restored and should be fully functional now. There will be a partial outage for ACCESS XDMoD today from approximately 10:00 EDT on 4/29 until 17:00 EDT on 5/01. This will only impact viewing job-level performance data of the Single Job Viewer. Other features of this service will still be fully functional during this time.

Posted: March 20, 2026 • Author: Conner Saeli

ACES Partial Availability, April 22-26

Published

Infrastructure News Type: Reconfiguration

Affected Infrastructure: aces.tamu.access-ci.org

Start Date: April 22, 2024, 2:00 p.m.

End Date: April 26, 2024, 11:00 p.m.

During April 22-26, 60 PVCs and 30 H100s will be unavailable while four racks are migrated from PCIe Gen4 fabric hardware to PCIe Gen5 fabric hardware.

Posted: March 20, 2026 • Author: Francis Dang

Ookami has two new NVIDIA Grace CPUs (144 cores each)

Published

Infrastructure News Type: Reconfiguration

Affected Infrastructure: ookami.sbu.access-ci.org

Start Date: April 15, 2024, 5:00 a.m.

End Date: May 31, 2024, 11:00 p.m.

The Ookami team is pleased to announce the addition of two NVIDIA Grace superchips (https://www.nvidia.com/en-us/data-center/grace-cpu-superchip/) to Ookami (CPU only). These new nodes with 144 cores each are now available for your testing and experimental projects. Read more (https://www.stonybrook.edu/commcms/ookami/support/faq/NVIDIA%20Grace%20CPUs.php)

Posted: March 20, 2026 • Author: Eva Siegmann

Anvil Cluster Maintenance

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: anvil.purdue.access-ci.org, anvil-gpu.purdue.access-ci.org

Start Date: March 12, 2024, 12:00 p.m.

End Date: March 13, 2024, 12:05 a.m.

Update as of March 12th 2024 8:05pm EDT The maintenance is now concluded and Anvil has been returned to service. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/open-a-ticket if you have any questions. Update as of March 12th 2024 5:42pm EDT Engineers are experiencing multiple outages at this time due to power disruption. Due to this, the scheduled Anvil cluster maintenance is being extended until Tuesday, March 12th at 10pm. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/open-a-ticket if you have any questions. Original News The Anvil cluster is unavailable beginning Tuesday, March 12th 2024 at 8:00am for a scheduled maintenance. It will return to full production by Tuesday, March 12th at 5pm. During this time, Anvil will have rack and power maintenance performed. Any jobs requesting a walltime which would take them past Tuesday, March 12th, 2024 at 8:00am will not start and will remain in the queue until after the maintenance is completed. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/open-a-ticket if you have any questions.

Posted: March 20, 2026 • Author: Ruyi Li

Kerberos Outage

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: kerberos.access-ci.org

Start Date: March 5, 2024, 7:00 p.m.

End Date: March 5, 2024, 10:00 p.m.

The Master Kerberos KDC will be upgraded to RHEL8. DUring this time users will not be able to create account or change passwords. Authentication should not be affected

Posted: March 20, 2026 • Author: Jacob Gallion

SDSC Expanse maintenance, 8AM-4PM (PT), Monday, 02/12/2024 [Completed]

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org

Start Date: February 12, 2024, 4:00 p.m.

End Date: February 13, 2024, 12:00 a.m.

>>> Update The Expanse maintenance is complete and the reservation has been released and jobs are running. Slurm has been updated to version 23.02.7. Please contact us via the ACCESS ticketing system if you have any questions. >>> Original message We will have a maintenance period on Expanse 8AM-5PM (PT), Feb 12, 2024. During this maintenance, we will be updating the Slurm scheduler. We have a reservation in place to prevent jobs from running during this period. The "squeue" output will show "ReqNodeNotAvail, Reserved for maintenance" for jobs that do not fit in the time period before the maintenance begins. These jobs will run after we release the maintenance reservation.

Posted: March 20, 2026 • Author: Mahidhar Tatineni