État

État des systèmes

Santé en direct de l'API, du tableau de bord, du traitement des paiements et du pool d'opérateurs.

Tous les systèmes sont opérationnels

Aucun incident en cours. L'API, le tableau de bord, le traitement des paiements, la livraison des SMS et celle des webhooks répondent normalement.

Vérifié 12:51 UTC

Disponibilité globale, 90 derniers jours

99.96 %

Composants

  • API REST
    Opérationnel
    Il y a 90 joursDisponibilité sur 90 jours · 99.98 %Aujourd'hui

    Temps de réponse médian de l'API 74 ms

  • Tableau de bord
    Opérationnel
    Il y a 90 joursDisponibilité sur 90 jours · 99.99 %Aujourd'hui

    Temps de réponse médian de l'API 168 ms

  • Paiements (OxaPay)
    Opérationnel
    Il y a 90 joursDisponibilité sur 90 jours · 99.94 %Aujourd'hui

    Temps de réponse médian de l'API 240 ms

  • Réception des SMS
    Opérationnel
    Il y a 90 joursDisponibilité sur 90 jours · 99.91 %Aujourd'hui
  • Envoi des webhooks
    Opérationnel
    Il y a 90 joursDisponibilité sur 90 jours · 99.96 %Aujourd'hui

    Temps de réponse médian de l'API 96 ms

Toutes les heures sont en UTC. La disponibilité est mesurée par composant, pondérée par la part de trafic affectée.

Historique des incidents

Les fenêtres de maintenance sont annoncées ici au moins 48 heures à l'avance.

Demander au support de vous ajouter

juillet 2026

DégradéRésolu

Webhook deliveries failing to endpoints behind one certificate chain

Impact
Envoi des webhooks
Durée
08:52–11:18 UTC · 2 h 26 min
  1. 08:52 UTC

    Webhook deliveries to a subset of endpoints are failing with a TLS verification error. Events are queued and will be retried, so nothing is being dropped. Investigating.

  2. 09:14 UTC

    A base image update on the delivery workers shipped a trust store that no longer carries a cross-signed intermediate that several customer endpoints still serve. Around 4 % of registered endpoints are affected.

  3. 09:48 UTC

    We have pinned the previous trust store on the delivery workers and the queued deliveries are going out. The oldest queued event is 56 minutes old.

  4. 11:18 UTC

    Resolved. The queue is empty and every event from the window was delivered well inside the 24-hour retry budget. The trust store is now pinned explicitly and updated deliberately instead of being inherited from the base image.

Panne partielleRésolu

Balance credits delayed by a payment callback backlog

Impact
Paiements (OxaPay)
Durée
11:06–12:24 UTC · 1 h 18 min
  1. 11:06 UTC

    Top-up credits are arriving late. Payments are being received and recorded, and no funds are at risk, but balances are updating with a delay. Investigating.

  2. 11:21 UTC

    OxaPay moved their callback traffic to a new source range this morning. Our allowlist rejected it, so callbacks were retried instead of accepted. About 340 callbacks are pending.

  3. 11:44 UTC

    The new range is allowlisted and the retries are being accepted. The backlog is draining oldest first at roughly 90 callbacks a minute.

  4. 12:24 UTC

    Resolved. Every pending credit has been applied and the longest delay was 78 minutes. We have asked to be notified of source-range changes in advance, and we now alert on a callback rejection rate above 1 % rather than only on queue depth.

juin 2026

MaintenanceRésolu

Singapore maintenance window overran

Impact
API REST, Tableau de bord
Durée
18:00–18:47 UTC · 47 min
  1. 18:00 UTC

    Maintenance window open, as announced on 2026-06-05. The Singapore entry point is draining and requests are being served from Frankfurt for the duration.

  2. 18:12 UTC

    The kernel upgrade is finished but one node is not rejoining the load balancer. Singapore stays drained while we look at it. Frankfurt is serving all traffic and latency from Asia is elevated to roughly 300 ms.

  3. 18:31 UTC

    The node was presenting a certificate issued before the upgrade. It has been reissued and the node is back in rotation.

  4. 18:47 UTC

    Resolved. Singapore is serving again at normal latency. We announced a 20-minute window and the drain lasted 47 minutes, which we should have said here sooner than we did.

Panne partielleRésolu

Inbound SMS dropped on an Indonesian operator

Impact
Réception des SMS, Envoi des webhooks
Durée
05:44–12:26 UTC · 6 h 42 min
  1. 05:44 UTC

    Activations on Indonesian numbers have been expiring without a code at a much higher rate than normal since about 05:10. Purchases and the rest of the catalogue are unaffected.

  2. 06:30 UTC

    The operator changed the route for inbound international SMS overnight without notice. Messages reach their network and are not forwarded to us. The affected ranges are out of routing as of 06:24.

  3. 07:05 UTC

    Indonesia is now served entirely by the two remaining operators. Success rate is 88 % against 93 % normally, and stock is about a third lower. Webhook volume for Indonesia is down accordingly.

  4. 09:40 UTC

    The operator has acknowledged the route change and is reverting it. We will not return the ranges to routing until our own probes have delivered cleanly for two hours.

  5. 12:26 UTC

    Resolved. The ranges are back in routing and delivering normally. 1,946 activations expired without a code during the incident; every one of them was refunded in full, automatically, with no ticket required.

mars 2026

DégradéRésolu

Rate limiter rejecting requests that were inside their budget

Impact
API REST
Durée
07:29–09:11 UTC · 1 h 42 min
  1. 07:29 UTC

    A number of API keys are receiving 429 responses while comfortably inside their per-minute budget. Investigating.

  2. 07:48 UTC

    A counter shard lost its expiry during a cache node replacement overnight, so counts from the previous hour were never cleared for keys hashed onto that shard. Roughly 6 % of keys are affected.

  3. 08:15 UTC

    The affected shard has been flushed and the keys on it are being served normally. We are checking the remaining shards for the same condition.

  4. 09:11 UTC

    Resolved. Counters now carry an absolute expiry written at creation instead of one set after the first increment, so a node replacement cannot leave a counter without one. No account was charged for a request that was rejected.

janvier 2026

Panne partielleRésolu

Top-up invoice creation failing upstream

Impact
Paiements (OxaPay)
Durée
10:22–13:05 UTC · 2 h 43 min
  1. 10:22 UTC

    Creating a top-up invoice is failing with an upstream error for most currencies. Invoices already created are unaffected and payments already sent are being credited normally.

  2. 10:40 UTC

    OxaPay has confirmed an incident on their invoice API. Purchases, activations, rentals and the rest of the API are not affected; only creating a new top-up is.

  3. 11:35 UTC

    Partial recovery upstream. USDT-TRC20 and TON invoices are being created again. BTC and ETH still fail intermittently.

  4. 12:41 UTC

    All currencies are creating invoices again. We are keeping this open while we watch the error rate.

  5. 13:05 UTC

    Resolved. 214 invoice attempts failed during the window and none of them took a payment. The top-up page now names the specific currency that is unavailable instead of failing generically.

novembre 2025

Panne partielleRésolu

Dashboard failed to load after a deploy

Impact
Tableau de bord
Durée
15:38–16:19 UTC · 41 min
  1. 15:38 UTC

    The dashboard is failing to load for some visitors with a chunk loading error. The API is unaffected, so automations are still running normally.

  2. 15:47 UTC

    A deploy at 15:31 invalidated the asset manifest while open sessions were still holding the previous one. Anyone who had the dashboard open before the deploy is affected; a hard reload works around it.

  3. 16:02 UTC

    The previous asset bundle has been re-published alongside the new one so both manifests resolve. Error reports have stopped.

  4. 16:19 UTC

    Resolved. Old asset bundles are now retained for 24 hours after a deploy, and the client reloads itself once when it sees a manifest it does not recognise.

octobre 2025

Panne partielleRésolu

Upstream operator degraded in Brazil

Impact
Réception des SMS
Durée
09:15–17:48 UTC · 8 h 33 min
  1. 09:15 UTC

    Success rates on Brazilian numbers have fallen from 91 % to about 44 % since 08:30. Purchases still succeed, but a large share of activations are expiring without a code.

  2. 09:52 UTC

    One operator range is being rejected by several services at once, which normally means the range has been flagged rather than that delivery is broken. Its pool score is set to zero so routing avoids it.

  3. 11:10 UTC

    Brazilian success rate is back to 86 % on the remaining operators. Stock for Brazil is roughly 40 % below normal while the flagged range is excluded.

  4. 14:30 UTC

    The operator confirms the range was flagged by a downstream aggregator and is migrating the block. We are keeping it out of routing until our own probes show it recovering.

  5. 17:48 UTC

    Resolved. Brazil is at 90 % success on a reduced pool. Every activation that expired during the incident was refunded automatically. The flagged range stays out of routing until the migration is finished.

août 2025

Panne majeureRésolu

Primary database failover

Impact
API REST, Tableau de bord, Paiements (OxaPay), Envoi des webhooks
Durée
04:07–04:52 UTC · 45 min
  1. 04:07 UTC

    The API is returning 500 for most requests and the dashboard will not load. The primary database is not responding. Investigating at the highest priority.

  2. 04:14 UTC

    The primary lost its storage volume. Automatic failover did not trigger because the primary was still answering health checks on its keepalive port. We are promoting the standby manually.

  3. 04:26 UTC

    The standby is promoted and the API is answering again. Error rate is back to baseline. We are now verifying whether any committed purchase was lost in the promotion.

  4. 04:52 UTC

    Resolved. 26 seconds of writes were lost in the promotion, covering 11 purchases. All 11 were refunded in full and the accounts were emailed individually. Health checks now run a real query instead of probing the keepalive port, and failover is exercised weekly.

juin 2025

MaintenanceRésolu

Scheduled maintenance: PostgreSQL major version upgrade

Impact
API REST, Tableau de bord, Paiements (OxaPay), Envoi des webhooks
Durée
02:00–02:41 UTC · 41 min
  1. 02:00 UTC

    Maintenance window open, as announced on 2025-06-10. Writes return 503 with a Retry-After header; reads and long-polls continue to be served from the replica.

  2. 02:18 UTC

    The upgrade is complete and the primary is accepting writes. We are comparing query plans against the pre-upgrade baseline before calling this done.

  3. 02:41 UTC

    Resolved. Total write unavailability was 18 minutes, inside the 45-minute window we announced. Activations that were open during the window were extended by 18 minutes so that no one lost part of an activation to the maintenance.

avril 2025

Panne partielleRésolu

Balance credits delayed by a stalled payment consumer

Impact
Paiements (OxaPay)
Durée
18:41–21:05 UTC · 2 h 24 min
  1. 18:41 UTC

    Top-ups paid in the last half hour are not appearing on balances. Payments are being received and recorded; it is the credit step that is backed up. No funds are at risk.

  2. 19:02 UTC

    The consumer that applies OxaPay callbacks to balances stopped acknowledging messages after a deploy at 18:10, and the queue has grown to about 400 callbacks. We are rolling that deploy back.

  3. 19:35 UTC

    The rollback is live and the backlog is draining at roughly 60 callbacks a minute. The oldest pending credit is 51 minutes old.

  4. 20:20 UTC

    Backlog cleared. Every payment received during the incident has been credited, including four invoices that expired while their callback was waiting.

  5. 21:05 UTC

    Resolved. The deploy removed the acknowledgement path for a callback that arrives for an invoice already in a terminal state, so those callbacks were retried forever and blocked the queue behind them. They are now acknowledged and logged.

mars 2025

DégradéRésolu

Elevated SMS delivery times on Indian ranges

Impact
Réception des SMS
Durée
06:52–10:34 UTC · 3 h 42 min
  1. 06:52 UTC

    Delivery times for Indian numbers are running above 60 seconds against a normal median of 9. Numbers are still being issued and the refund window is unchanged. Investigating.

  2. 07:20 UTC

    One of our two Indian upstreams is queueing inbound SMS at its own gateway. India is being routed to the second provider while we wait on their side.

  3. 08:05 UTC

    The shift is complete and new Indian activations are delivering at a median of 14 seconds. Activations bought before 07:20 that are still open will either deliver or auto-refund at the end of their window.

  4. 09:40 UTC

    The upstream has drained its queue and reports a full disk on a gateway node as the cause. We are keeping traffic on the second provider until their median has been stable for an hour.

  5. 10:34 UTC

    Resolved. India is back on both providers at a median of 9.2 seconds. 812 activations expired without a code during the incident and were refunded automatically.

Être notifié

Changements d'état par e-mail ou par webhook.

S'abonner au flux RSS