System and Infrastructure Status News

ACES Lustre filesystem issues

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: aces.tamu.access-ci.org

Start Date: July 1, 2025, 8:00 p.m.

End Date: July 2, 2025, 12:05 a.m.

We are currently seeing degradation on one of the Lustre storage servers. This is leading to slow filesystem access and impacting the responsiveness of the Slurm job scheduler. We will update once the issue is resolved. The Lustre filesystem has been recovered. We monitoring the storage for any further issues.

Posted: March 20, 2026 • Author: Francis Dang

SDSC Expanse Lustre filesystem issues

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org

Start Date: June 18, 2025, 8:00 a.m.

End Date: June 18, 2025, 4:00 p.m.

Dear Expanse User, We are currently seeing connectivity issues to one of the Lustre filesystem object storage servers (OSSs). This is leading to timeouts and access issues for files that are striped onto this OSS. We will update once the issue is resolved. Thanks SDSC User Services Staff

Posted: March 20, 2026 • Author: Mahidhar Tatineni

Anvil power outage

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: anvil.purdue.access-ci.org, anvil-gpu.purdue.access-ci.org

Start Date: June 14, 2025, 4:30 p.m.

End Date: June 14, 2025, 7:40 p.m.

Update 3:40PM, EDT June 14, 2025 Our engineers have brought Anvil back to full service. If you have any questions about this outage, please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket. Thank you. Original Post: Shortly after 12:30pm June 14, 2025, we had a major power outage at our data center at Purdue. Anvil has been impacted and will be offline until power is resumed. Our engineering team is working closely with campus power engineers to bring the power and Anvil back. We apologize for any inconvenience it might have caused. There is no ETA yet, but we will provide an update as soon as we have one. Please submit a ticket through ACCESS Help Desk at https://support.access-ci.org/help-ticket if you have any questions.

Posted: March 20, 2026 • Author: Guangzhen Jin

[Resolved] Anvil Network Interruption

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: anvil.purdue.access-ci.org, anvil-gpu.purdue.access-ci.org

Start Date: June 10, 2025, 2:00 p.m.

End Date: June 10, 2025, 11:30 p.m.

Starting at 10:00 AM EST, Anvil began experiencing a network interruption. This may impact access to services such as Open OnDemand and Remote Desktop. Additional impacts are still being assessed. At this time, we do not have an estimated time for service restoration. We will provide updates as more information becomes available. Thank you for your patience and understanding. Updates: The issue was resolved at 7:30 PM EST on June 10. All services should now be functioning normally. If you still experience any problem, please feel free to reach out Anvil support team through ACCESS Help Desk (https://support.access-ci.org/help-ticket).

Posted: March 20, 2026 • Author: Nannan Shan

ACES Maintenance - June 4-5

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: aces.tamu.access-ci.org

Start Date: June 4, 2025, 2:00 p.m.

End Date: June 5, 2025, 5:00 p.m.

The ACES cluster will be unavailable during maintenance from 9am to 8pm CDT on Wednesday June 4. A reservation is in place to prevent jobs from running past the start time of the maintenance period. The maintenance is currently extended to 12pm CDT Thursday.

Posted: March 20, 2026 • Author: Francis Dang

SDSC Machine room power outage [Expanse, Voyager returned to production]

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org

Start Date: May 31, 2025, 9:30 p.m.

End Date: June 2, 2025, 6:30 a.m.

>>>>> Update 1 Dear Expanse User, Expanse was put back in production after recovery from the power outage we had in the machine room (and UCSD wide). The machine is available for use and running jobs. Voyager has also been brought back into production. Thanks SDSC User Services Staff >>>>> There was a power outage at UCSD that impacted the SDSC machine room. The systems at SDSC (Expanse, Voyager) are currently down. We will update once they are brought up and accessible.

Posted: March 20, 2026 • Author: Mahidhar Tatineni

Upcoming Changes to the Ticketing System Portal Forms

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: tickets.access-ci.org

Start Date: May 29, 2025, 2:00 p.m.

End Date: June 5, 2025, 2:00 p.m.

Hi Everyone, The ACCESS Operations team is working on optimizing the ticketing system by streamlining the dropdown options available on portal forms. These improvements are intended to enhance the user experience and make ticket categorization more intuitive. We are planning to roll out these changes on June 6, and the implementation window will take place from May 29 to June 5. During this period, there may be minor outages or brief disruptions. Updated Operational Support Issues: - Allocations (includes AMIE, XRAS, etc.) - Support (includes OnDemand, Pegasus, Knowledge Base, Affinity Groups, Events, Announcements, Ask.CI, etc.) - Security and Authentication (includes IAM, policies, etc.) - Networking and Data Transfer (includes SSL, DNS, CONECTnet, etc.) - Operations Infrastructure (includes monitoring, logging, GitHub, etc. - Operations Software and Online Services (includes CiDeR, portal, dashboard, API, etc.) - Metrics - ACCESS Communications and Collaboration Tools - Ticket System - Resource Integration - Some Other Question Please note that minor updates to queues and watcher groups will also be included as part of this change. We appreciate your patience and understanding as we work to improve the system. We'll keep you informed with any further updates. Thanks & Regards, Dinuka (On Behalf of ACCESS Operations)

Posted: March 20, 2026 • Author: Dinuka De Silva

SDSC Expanse Lustre filesystem issues (update)

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org

Start Date: May 25, 2025, 10:45 p.m.

End Date: May 29, 2025, 12:00 a.m.

>>> Update #3 Dear Expanse User, We have remounted the Lustre filesystem on Expanse excluding the OST that is having problems. This should prevent the hung tasks that were causing higher loads on the login nodes. We continue to work on the OST with issues and will update once it is returned to service. In the interim any old files that were striped onto the OST will fail on reads. New I/O to the filesystem will target healthy OSTs. Thanks SDSC User Services Staff >>> Update #2 Dear Expanse User, We are still working on the one OST from the Expanse Lustre filesystem that is failing to mount. This is making all files/directories that are striped onto this OST unavailable. Please note that this will also cause full file listings to hang so please refrain from doing full metadata listings on Lustre directories. All other OSTs on the filesystem are usable and new I/O will automatically avoid the problem OST. We will update once the OST with issues is restored. Thanks SDSC User Services Staff >>> Update Dear Expanse User, We brought the two OSSs online last night but there is still one storage target on one of them that needs more work to recover. We are continuing to look at the issue and will update again once the filesystem is back. Thanks SDSC User Services >>> Dear Expanse User, We are currently seeing problems with two object storage servers (OSSs) that are part of the Expanse Lustre filesystem. This is causing access issues on files that are striped on these servers. Please refrain from doing full metadata listings on Lustre directories as chances are you will access a file that is on one of the OSSs and the commands might hang. We are working on resolving the problem and will update once the OSSs are back in service. Thanks SDSC User Services Staff

Posted: March 20, 2026 • Author: Mahidhar Tatineni

ACES Partial Maintenance - May 7

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: aces.tamu.access-ci.org

Start Date: May 7, 2025, 1:00 p.m.

End Date: May 7, 2025, 10:00 p.m.

The nodes and devices attached to the composability fabrics in ACES Racks 2 and 3 will be offline for maintenance on Wednesday, May 7 from 8:00 AM CDT to 5:00 PM CDT.

Posted: March 20, 2026 • Author: Francis Dang

Launch Maintenance - May 7

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: launch.tamu.access-ci.org

Start Date: May 7, 2025, 1:00 p.m.

End Date: May 7, 2025, 10:00 p.m.

The TAMU Launch cluster will be down for maintenance on Wednesday, May 7 from 8:00AM CDT to 5:00PM CDT.

Posted: March 20, 2026 • Author: Francis Dang

Ticketing System Automation Rule execution is delayed

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: tickets.access-ci.org

Start Date: April 23, 2025, 3:00 p.m.

End Date: April 24, 2025, 7:46 p.m.

Hi Everyone, Atlassian is investigating the cause of the rule execution delay. This might delay the communication of the tickets for you. We will keep you posted when we have an update. Meanwhile, your patience and understanding are appreciated. https://jira-service-management.status.atlassian.com/incidents/4717qqxyt0nk Thanks

Posted: March 20, 2026 • Author: Dinuka De Silva

SDSC Expanse: Scheduler issues [Resolved]

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org

Start Date: April 9, 2025, 11:30 p.m.

End Date: April 10, 2025, 4:15 a.m.

The Expanse scheduler issue has been fixed and job submissions and queue commands are working now. Thanks SDSC User Services Staff ----- Dear Expanse User, We are currently seeing issues with the Expanse Slurm scheduler and troubleshooting the problem. At present new job submissions and Slurm commands are failing. We will update once the issue is resolved. Thanks SDSC User Services

Posted: March 20, 2026 • Author: Mahidhar Tatineni

Update to Production Deployment of RP SSH Pubkey Service

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: registry.access-ci.org

Start Date: April 9, 2025, 2:00 p.m.

End Date: April 9, 2025, 3:00 p.m.

On April 9, 2025, starting at 9:00am CDT (10:00am EDT) ACCESS Operations will be updating the production deployment of the SSH Public Key retrieval service for ACCESS RPs. This will require a very brief downtime for the existing service, during which queries from RP registered to use the service may be interrupted. There will be no interruption to the ACCESS User Identity and OAuth Client Registry for this update.

Posted: March 20, 2026 • Author: Derek Simmel

502 Bad Gateway Error Affecting JSM Customer Portal in Some Regions

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: tickets.access-ci.org

Start Date: March 25, 2025, 6:29 a.m.

End Date: March 25, 2025, 6:49 a.m.

Users in some regions are currently experiencing a 502 Bad Gateway error when trying to access the JSM ticketing system Customer Portal. This issue is impacting both customers and agents in those regions. We are waiting for an update from Atlassian support and will provide further information as soon as we hear from them. https://jira-service-management.status.atlassian.com/incidents/vy4wm8b5v0yh

Posted: March 20, 2026 • Author: Dinuka De Silva

Some errors in JSM automation executions

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: tickets.access-ci.org

Start Date: March 24, 2025, 12:39 a.m.

End Date: March 24, 2025, 4:56 a.m.

Hi Everyone, We are encountering some errors in the automation executions and as a result, the notifications for new tickets or ticket edits may not go out. Other than that, there's no impact on managing tickets or queues. Please bear with us while we are troubleshooting this with Atlassian support. We will keep you posted as we know. Thank you

Posted: March 20, 2026 • Author: Dinuka De Silva

Unplanned outage to xdmod.access-ci.org and metrics.access-ci.org [RESOLVED]

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: xdmod.access-ci.org

Start Date: March 21, 2025, 10:28 a.m.

End Date: March 21, 2025, 2:55 p.m.

We are investigating an issue with connectivity to xdmod.access-ci.org and metrics.access-ci.org. Resolved 9:56 AM EDT

Posted: March 20, 2026 • Author: Joseph White

Kerberos Replica DNS Work

Published

Infrastructure News Type: Reconfiguration

Affected Infrastructure: kerberos.access-ci.org

Start Date: March 18, 2025, 6:00 p.m.

End Date: March 18, 2025, 7:00 p.m.

There will be DNS work to transition to the new Kerberos replica server hosted at PSC. No service outage expected as hosts should failover to the other replicas.

Posted: March 20, 2026 • Author: Jacob Gallion

SDSC Expanse: Lustre filesystem back in production use (03/24/2025)

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: expanse.sdsc.access-ci.org, expanse-gpu.sdsc.access-ci.org, expanse-ps.sdsc.access-ci.org

Start Date: March 17, 2025, 4:00 p.m.

End Date: March 25, 2025, 12:00 a.m.

Dear Expanse User, The Expanse Lustre filesystem issues have been resolved and the filesystem is back in production use. Thank you for your patience through the long outage. We will continue to monitor the filesystem and follow up as needed. Users are also reminded that both /expanse/lustre/scratch and /expanse/lustre/projects are not backed up. So please make offsite copies of anything critical in those locations. Thanks SDSC User Services Staff ------------ Dear Expanse User, We are continuing to work on the Expanse Lustre filesystem. The initial problem was due to a hardware issue with one of the metadata server drives. The drive was replaced but problems persist with mounting of storage targets due to a software bug being triggered in the Lustre filesystem. Unfortunately since the metadata server controls the entire filesystem, the /expanse/lustre/scratch and /expanse/lustre/projects directories will continue to be unavailable. We are sorry for the impact this is causing and will keep users posted about any new developments. Thanks SDSC User Services Staff --------- Dear Expanse User, We are continuing to work on the metadata server problem on the Expanse Lustre filesystem. Unfortunately the outage is going to go longer and we will update once we have more information. Thanks SDSC User Services Staff ----------------------------------------- Dear Expanse User, We are continuing to work on the Lustre filesystem on Expanse. The problem is going to take much longer than anticipated to resolve and likely the earliest we can recover is tomorrow (03/18/2025). We recognize that a lot of Expanse users do not use the lustre directories and to enable them to run we will release the reservation. We have held current jobs that are clearly using lustre (e.g. if they specified it in the output path or working directory path). However, if you have jobs that use Lustre without specifying the need through a constraint, the jobs will fail. We strongly recommend all jobs needing Lustre include the following line: #SBATCH --constraint="lustre" Please see more details in our user guide under the "SUBMITTING JOBS USING LUSTRE" subsection in the storage section (https://www.sdsc.edu/systems/expanse/user_guide.html#narrow-wysiwyg-10). We also want to remind users that the home and NFS directories are limited in performance and scaling so please don't submit intensive jobs that need the Lustre filesystem from there. We will keep you posted on the Lustre issue status tomorrow. Thanks SDSC User Services Staff ----------------------------------------

Posted: March 20, 2026 • Author: Mahidhar Tatineni

Service Slowness in Multiple Products (Jira, Jira Service Management, and Confluence)

Published

Infrastructure News Type: Outage Full

Affected Infrastructure: tickets.access-ci.org

Start Date: March 17, 2025, 3:30 p.m.

End Date: March 18, 2025, 4:31 a.m.

This may have delayed some of the automation executions up to one day. But, as of now, there's no any other impact identified. https://jira-service-management.status.atlassian.com/incidents/hbhx13fzvdgz#

Posted: March 20, 2026 • Author: Dinuka De Silva

ACCESS XDMoD Partial Downtime

Published

Infrastructure News Type: Outage Partial

Affected Infrastructure: xdmod.access-ci.org

Start Date: March 17, 2025, 2:00 p.m.

End Date: March 18, 2025, 2:00 p.m.

ACCESS XDMoD will be upgraded to version 11.0.1 on Monday, March 17 at approximately 10:00 EDT. Various data in ACCESS XDMoD may be unavailable during the upgrade. Service is expected to be fully restored within 24 hours. Once the upgrade is started, release notes will be available at https://xdmod.access-ci.org/#main_tab_panel:about_xdmod?Release%20Notes

Posted: March 20, 2026 • Author: Aaron Weeden