HARICA nos informa de lo siguiente:
Subject: HARICA update on recent mass-replacement of TLS Certificates
Dear TCS-PMA Members,
Once again, we wish to express our sincere regret for the inconvenience caused by the two recent incidents. We want to be absolutely clear that in both cases, the affected certificates remained fully trustworthy; there was no compromise of security, private key material, or validation integrity at any point. Rather, both incidents stemmed from technical discrepancies between the strict language of our CP/CPS and the implementation of industry best practices.
Please find below an update on the current status of the certificate reissuance process:
- ACME ARI Reissuance: All ACME ARI endpoints have been updated, and all affected certificates are expected to be automatically replaced before the revocation deadline. We are closely monitoring ACME issuance, and current rates are high and stable. While minor delays and timeouts have occasionally occurred during peak loads on the CertManager application, our infrastructure administrators are monitoring the systems 24/7 to resolve any issues immediately. We expect all ACME ARI-capable clients to automatically complete certificate replacements by Friday at 23:59:59 UTC.
- Non-ACME (Reusable Validation): For non-ACME certificates with valid, reusable Domain (DCV) and Identity Validation, we have successfully reissued the certificates and emailed the Subscribers with instructions to retrieve them. Out of the total population of approximately 235,000 affected TLS certificates (across all operations, not just TCS), approximately 1,500 certificates could not be automatically replaced due to CAA, MPIC, or other technical restrictions. Our support team is actively working directly with the relevant administrators to coordinate resolutions for these specific cases.
- Non-ACME (Expired Validation): For non-ACME certificates with expired DCV or Identity Validation data, our support team has contacted the respective administrators to request the submission of new certificate applications.
- Overall Progress: As of 14:30 CEST today (2026-07-23), we verified that 71.36% of the affected Domain Names have already been secured with new certificates. We expect this percentage to rise steadily over the next two days as ACME ARI continues to schedule automated reissuances through Friday night.
Challenges Encountered During Mass-Replacement
We observed a few behaviors that complicated the initial replacement efforts:
- Unexpected System Load: Many Subscribers acted immediately upon receiving the initial notification on Tuesday, manually replacing certificates despite the email indicating that HARICA would handle the process smoothly. This concurrent manual activity spiked system load precisely as we were staging the automated mass-replacement. For example, some ACME clients forced immediate updates, bypassing the ARI scheduling mechanisms designed to safely distribute server load over time.
- Duplicate Issuance: Several Subscribers utilized automated API calls to request new certificates for those slated for replacement. This resulted in duplicate issuance—one triggered manually by the Subscriber and another processed simultaneously by HARICA's mass-replacement jobs.
Immediate Mitigations Implemented
During the course of this mass-replacement event, we took the following emergency steps to stabilize and enhance service performance:
- Allocated additional compute resources across all affected CA, RA, and Database servers.
- Added new database indexes to optimize and accelerate query speeds.
- Deployed an additional server node for CertManager to increase availability and throughput.
- Adjusted throttling thresholds to better accommodate high volumes of outgoing emails and server connections.
Planned Infrastructure Improvements
In the near future, we will be implementing the following permanent enhancements:
- Introducing rate limiting (throttling) for REST API requests to prevent future resource exhaustion.
- Adding permanent CertManager nodes to the RA cluster.
- Deploying additional database nodes for greater resilience.
- Refactoring code to further optimize the efficiency of database queries and statements.
Root Cause Analysis (RCA)
Regarding the Root Cause Analysis for these two incidents, our investigations are actively underway. We will publish the comprehensive reports in the corresponding public Bugzilla trackers:
- https://bugzilla.mozilla.org/show_bug.cgi?id=2055551
- https://bugzilla.mozilla.org/show_bug.cgi?id=2056668
We wish to reiterate that we consider the GÉANT community and the TCS project to be of the utmost importance. Our team is working diligently and tirelessly to resolve systemic issues and address any infrastructural weaknesses identified during this process. Furthermore, we have scheduled a wide range of improvements and feature updates for the coming months to enhance the service overall.
We deeply appreciate your continued support and understanding. We would be happy to discuss this in greater detail via a teleconference next week, or at your earliest convenience.