System and Infrastructure Status News
FASTER Maintenance, March 10-13
PublishedInfrastructure News Type: Outage Full
Affected Infrastructure: faster.tamu.access-ci.org
Start Date: March 10, 2025, 2:00 p.m.
End Date: March 15, 2025, 5:00 p.m.
The FASTER cluster will be unavailable from 9am March 10 to 8pm March 13 for usual OS maintenance, Lustre storage maintenance, and Liqid fabric composability maintenance.
Posted: March 20, 2026 • Author: Francis Dang
Connection Errors for Jira Service Management in some regions
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: tickets.access-ci.org
Start Date: March 6, 2025, 5:27 p.m.
End Date: March 7, 2025, 7:00 a.m.
This incident affects: Jira Service Management Web, Service Portal, Opsgenie Incident Flow, Opsgenie Alert Flow, Opsgenie Incident Flow, Opsgenie Alert Flow, Jira Service Management Email Requests, Authentication and User Management, Purchasing & Licensing, Signup, Automation for Jira, and Assist. https://jira-service-management.status.atlassian.com/incidents/hjh7ydq8jlj6
Posted: March 20, 2026 • Author: Dinuka De Silva
SDSC Expanse: Upcoming change to require two-factor authentication on Expanse (Reminder)
PublishedInfrastructure News Type: Reconfiguration
Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org
Start Date: February 24, 2025, 5:00 p.m.
End Date: July 7, 2025, 4:00 p.m.
Dear Expanse User, Starting Feb 24, 2025, two-factor authentication (2FA) will be required on Expanse. To avoid losing access, set up 2FA with Google Authenticator before this date. Instructions are in the Expanse user guide (https://www.sdsc.edu/systems/expanse/user_guide.html) in the system access section. For role accounts needing programmatic access, Science Gateway PIs will be contacted with further details. Other users with automation needs should contact SDSC user support via ticketing before Feb 24, 2025. SDSC User Services Staff
Posted: March 20, 2026 • Author: Mahidhar Tatineni
DUO authentication issues for phone-call-based authentication leading to temporary lockouts.
PublishedInfrastructure News Type: Degraded
Affected Infrastructure: duo.access-ci.org
Start Date: February 13, 2025, 6:48 p.m.
End Date: February 14, 2025, 1:43 a.m.
DUO has reported issues with phone-call-based authentication on their systems today, leading to some users being locked out for several hours for too many authentication failures. We do NOT recommend that ACCESS users employ phone-call-based authentication, as it is often unreliable, and prone to being abused. We strongly recommend that ACCESS users configure their DUO authentication to use the DUO App on a mobile device, and to use Push authentication. Other authentication methods, including Passkey or token-based authentication are known to work well. For more information about setting up DUO, please consult the documentation at https://guide.duo.com/universal-prompt#add-or-manage-devices.
Posted: March 20, 2026 • Author: Derek Simmel
SDSC Expanse Lustre filesystem issues
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org
Start Date: February 9, 2025, 6:00 a.m.
End Date: February 9, 2025, 6:00 p.m.
Dear Expanse User, We are currently seeing high load on one of the metadata servers of the Expanse Lustre filesystem. This is leading to timeouts on access to some files and directories. We are looking into the problem and will update once it is resolved. SDSC User Services Staff
Posted: March 20, 2026 • Author: Mahidhar Tatineni
SDSC Expanse: Power infrastructure maintenance
PublishedInfrastructure News Type: Outage Full
Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org
Start Date: February 6, 2025, 5:00 a.m.
End Date: February 6, 2025, 4:00 p.m.
Dear Expanse User, We will be working on the SDSC machine room power infrastructure from 9PM (PT), Feb 5, 2025 to 8AM (PT) Feb 6, 2025. This will impact power to racks 1-8, 16, and 17 of Expanse. We have placed a maintenance reservation on these racks for the duration of the work. This will impact job wait times as we get closer to the maintenance period but the remaining resources of the system will available through the maintenance and logins, filesystem access, and job submissions are not expected to be impacted. Thanks SDSC User Services Staff
Posted: March 20, 2026 • Author: Mahidhar Tatineni
SDSC Expanse Scheduler issue [Resolved]
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org
Start Date: February 3, 2025, 1:00 a.m.
End Date: February 3, 2025, 5:00 a.m.
Update - the issue was resolved last night around 9PM (PT) Dear Expanse User We had a scheduler issue this evening that unfortunately led to jobs in the queue being lost. We apologize for the inconvenience this will cause as jobs will have to be resubmitted. We are looking into the issue and will update once the scheduler service is restored. Thanks SDSC User Services Staff
Posted: March 20, 2026 • Author: Mahidhar Tatineni
SDSC Expanse Lustre filesystem issues
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org
Start Date: January 28, 2025, 5:00 a.m.
End Date: January 28, 2025, 3:00 p.m.
Update: The Lustre MDS issue was resolved this morning and the filesystem access is back to normal. Dear Expanse User We are currently seeing issues with the Expanse Lustre filesystem. This is leading to very slow responses or timeouts on access. We will update once the problem is resolved. SDSC User Services Staff
Posted: March 20, 2026 • Author: Mahidhar Tatineni
Jira services are unavailable and have degraded performance in certain regions
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: tickets.access-ci.org
Start Date: January 23, 2025, 4:30 p.m.
End Date: January 24, 2025, 1:00 p.m.
We have been informed of the degraded performance of Jira Work Management, Jira Service Management, and Jira Cloud customers in certain regions. We will provide more details as soon as we have. This incident affects: Jira Service Management Web, Service Portal, Opsgenie Incident Flow, Opsgenie Alert Flow, Opsgenie Incident Flow, Opsgenie Alert Flow, Jira Service Management Email Requests, Authentication and User Management, Purchasing & Licensing, Signup, Automation for Jira, and Assist. https://jira-service-management.status.atlassian.com/incidents/4s58pz6sk3zj This has been resolved
Posted: March 20, 2026 • Author: Dinuka De Silva
Unscheduled Anvil Outage
PublishedInfrastructure News Type: Degraded
Affected Infrastructure: anvil.purdue.access-ci.org, anvil-gpu.purdue.access-ci.org
Start Date: January 21, 2025, 7:30 p.m.
End Date: January 21, 2025, 8:57 p.m.
Update: As of Tuesday, January 21st, 2025 at 3:57pm EST, this has been resolved and capacity has been restored. The Anvil cluster began experiencing issues with electrical power around 2:30 PM EST. RCAC engineers are working with Purdue electricians to safely restore power. Anvil is operating at reduced capacity while a handful of nodes were shut down as a precaution. If your jobs were running on these please resubmit. If you have any questions, please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket. We will provide an update by 5:00 PM.
Posted: March 20, 2026 • Author: Guangzhen Jin
Anvil Cluster Open Ondemand Maintenance - January 17, 2025
PublishedInfrastructure News Type: Reconfiguration
Affected Infrastructure: anvil.purdue.access-ci.org, anvil-gpu.purdue.access-ci.org
Start Date: January 17, 2025, 2:00 p.m.
End Date: January 17, 2025, 6:00 p.m.
Update: As of 12:00pm EDT. Jan 17, Anvil team has completed maintenance and returned the Open Ondemand service on Anvil cluster back to normal service. Please enjoy the new features on this dashboard and let us know if you notice any bugs or want more features by submitting a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket. Update: the maintenance has been postponed to Friday Januany 17, 2025 The Open Ondemand service for Anvil will be unavailable from Friday, January 17 at 9:00am EDT, 2025 to Friday, January 17 at 5:00pm EDT, 2025. During the maintenance, Anvil team will perform a reconfiguration to the Open Ondemand dashboard for Anvil which include a brand new design of the dashboard with new features listed below. What’s New on the dashboard? - Service Unit Balance and Usage: Monitor your allocation usages and remaining balance on Anvil. - Disk Usage: Monitor your storage utilization across Anvil's file systems. - Job Queue: View and manage your running and queued jobs on Anvil. - News Feed: Stay updated with the latest Anvil news and announcements. - Partition Status: Monitor the current state of partitions/queues on Anvil. - My Jobs Page: Re-designed page to show detailed job information for your jobs and jobs in your allocation(s) as well as job management. - Performance Metrics Page: Analyze your job performance and resource utilization patterns over time. What will impact you? - All Slurm jobs on Anvil (including jobs that have already submitted through Open Ondemand before this maintenance) will continue and NOT be impacted. - All functions including login to Open Ondemand will be unavailable during the maintenance. Anvil Open Ondemand service will return to full production by Friday, January 17 at 5:00pm EDT, 2025. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket (https://support.access-ci.org/help-ticket**) if you have any questions or suggestions.
Posted: March 20, 2026 • Author: Guangzhen Jin
ACES Maintenance - January 16
PublishedInfrastructure News Type: Outage Full
Affected Infrastructure: aces.tamu.access-ci.org
Start Date: January 16, 2025, 3:00 p.m.
End Date: January 17, 2025, 2:00 a.m.
The ACES cluster will be unavailable during maintenance from 9am to 8pm CST on Thursday January 16 A reservation is in place to prevent jobs from running past the start time of the maintenance period. After the maintenance has been completed, the maximum permitted time limit for jobs in the cpu queue will be reduced from 7 days to 3 days.
Posted: March 20, 2026 • Author: Francis Dang
Hive Gateway Retirement
PublishedInfrastructure News Type: Retirement
Affected Infrastructure: hive.gatech.access-ci.org
Start Date: January 13, 2025, 6:00 a.m.
End Date: Not Specified
Dear ACCESS Community, The Hive Gateway resource at Georgia Tech will enter retirement on January 13th, 2025. The original hardware has reached end-of-life after an additional year of performance post-award and can no longer be supported. The Gateway will remain accessible for existing users, in order to access any data on the system, but it will be impossible to run new jobs. We currently plan to fully turn off Gateway access for ACCESS accounts by March, 2025. Please feel free to reach out if you have any concerns via email at pace-support@oit.gatech.edu. Best, The PACE Team
Posted: March 20, 2026 • Author: John Coulter
ACCESS XDMoD Unplanned Partial Outage
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: xdmod.access-ci.org
Start Date: January 9, 2025, 6:00 p.m.
End Date: January 10, 2025, 11:00 p.m.
UPDATE 01/10 - The service has been fully restored at this time. There is a unplanned partial outage for ACCESS XDMoD from approximately 12:00 EDT on Thursday, January 9th until 17:00 EDT on Friday, January 10th. This is a partial outage that impacts viewing job performance data in the Single Job Viewer. An update will be posted once the issue is resolved.
Posted: March 20, 2026 • Author: Conner Saeli
Jetstream2 Planned Outage: January 6–9, 2025
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: jetstream2.indiana.access-ci.org, jetstream2-gpu.indiana.access-ci.org, jetstream2-lm.indiana.access-ci.org, jetstream2-storage.indiana.access-ci.org
Start Date: January 6, 2025, 3:00 p.m.
End Date: January 9, 2025, 3:00 p.m.
On Monday, January 6, 2025 at 9AM EST, Jetstream2 will begin a maintenance outage that will last through approximately 9AM EST on Thursday, January 9 (subject to change). This infrastructure maintenance is being done in conjunction with a third party vendor to update Jetstream2’s cooling system in order to accommodate a resource expansion. This maintenance outage will affect all primary Jetstream2 resources (CPU, GPU, Large Memory, and Storage) and user interfaces (Exosphere, CACAO). While existing instances at satellite regions will not be affected by the maintenance, the Exosphere and CACAO user interfaces will be inaccessible for all regions. When maintenance is complete, all instances will be returned to either a shelved or active state. If your instance was in an active, unshelved state, it will be brought back to that state. If it was shelved, it will remain shelved. NOTE: If your instance was in an errored, shutoff, or suspended state, it will be restored to an active state. We strongly advise all Jetstream2 users to review the states of their existing instances, as well as save and close their work prior to January 6. You can preserve your work by: - Safely shelving any active instances - Backing up essential data outside of Jetstream2 or creating images (https://docs.jetstream-cloud.org/general/instancemgt/#image) of your instances During the outage, please refer to the Jetstream2 status page (https://jetstream.status.io/) for the most up-to-date information. We appreciate your understanding and hope to mitigate any inconvenience this might cause. If you have any questions, please contact the Jetstream2 Support team at help@jetstream-cloud.org (mailto:help@jetstream-cloud.org?subject=).
Posted: March 20, 2026 • Author: Zachary Graber
Expanse Lustre filesystem issues [Resolved]
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org
Start Date: December 18, 2024, 10:30 a.m.
End Date: December 18, 2024, 7:00 p.m.
Update: The Lustre OSS with the problems was fixed and returned to service and the filesystem is accessible on Expanse now. Dear Expanse User We are currently seeing issues with the Expanse Lustre filesystem. This is leading to very slow responses or timeouts on access. We will update once the problem is resolved. SDSC User Services Staff
Posted: March 20, 2026 • Author: Mahidhar Tatineni
Expanse racks impacted by power maintenance
PublishedInfrastructure News Type: Outage Partial
Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org
Start Date: November 25, 2024, 4:30 p.m.
End Date: November 26, 2024, 1:00 a.m.
Dear Expanse User, A maintenance is ongoing to update part of the datacenter power infrastructure. This has impacted more Expanse racks than we originally anticipated (due to cooling considerations) and as a result some of the jobs running on Expanse were impacted. The jobs will get a NODE_FAIL error in Slurm so they will not be charged SUs. We will update once the maintenance is complete and the nodes are returned to service. In the interim, Expanse will have fewer available nodes than normal so wait times will likely increase today. We are sorry for the unexpected impact and please send us a ticket (either ACCESS or SDSC ticketing system) if you have any questions. Thanks SDSC User Services Staff
Posted: March 20, 2026 • Author: Mahidhar Tatineni
Discounted Exchange rate for ACCESS credits on Anvil CPU
PublishedInfrastructure News Type: Reconfiguration
Affected Infrastructure: anvil.purdue.access-ci.org
Start Date: November 19, 2024, 6:00 a.m.
End Date: March 31, 2025, 5:00 a.m.
Anvil CPU is now offering a discounted exchange rate for researchers with Explore, Discover, and Accelerate allocations. The revised exchange rate is now 1 ACCESS credit = 1 Anvil CPU service unit (core hour). This new rate will be applicable for all exchanges to Anvil CPU starting 11/19/2024. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket if you have any questions.
Posted: March 20, 2026 • Author: Guangzhen Jin
Anvil Cluster Maintenance, October 21-22
PublishedInfrastructure News Type: Outage Full
Affected Infrastructure: anvil.purdue.access-ci.org, anvil-gpu.purdue.access-ci.org
Start Date: October 21, 2024, 11:00 a.m.
End Date: October 22, 2024, 1:00 a.m.
UPDATE: October 21, 2024 09:00 PM EDT As of 9:00PM EDT, the maintenance work on Anvil has been completed and job scheduling has been resumed. If you encounter any issues post maintenance, please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket. Original Post: The Anvil system will be unavailable from Monday, October 21st at 7:00am to Tuesday, October 22nd at 6:00pm EDT, 2024 for scheduled maintenance. During the maintenance, we will perform cooling work for the system to integrate a new NSF funded AI resource into Anvil. Any Slurm jobs which request a walltime which would take them past Monday, October 21st at 7:00am EDT will not start and will remain in the queue until after the maintenance is completed. Anvil will return to full production by Tuesday, October 22nd at 6:00pm, EDT 2024. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket if you have any questions.
Posted: March 20, 2026 • Author: Guangzhen Jin
FASTER Maintenance, October 15
PublishedInfrastructure News Type: Outage Full
Affected Infrastructure: faster.tamu.access-ci.org
Start Date: October 15, 2024, 2:00 p.m.
End Date: October 16, 2024, 1:00 a.m.
The FASTER cluster will be unavailable during maintenance from 9am to 8pm CDT on Tuesday October 15. A reservation is in place to prevent jobs from running past the start time of the maintenance period.
Posted: March 20, 2026 • Author: Francis Dang