状态

系统状态

API、控制台、支付处理与运营商资源池的实时健康状况。

所有系统运行正常

当前没有正在进行的故障。API、控制台、支付处理、短信送达和 Webhook 投递均响应正常。

11:35 UTC 检查

整体可用率,最近 90 天

99.96 %

组件

  • REST API
    正常
    90 天前90 天可用率 · 99.98 %今天

    API 响应中位时间 74 ms

  • 控制台
    正常
    90 天前90 天可用率 · 99.99 %今天

    API 响应中位时间 168 ms

  • 支付(OxaPay)
    正常
    90 天前90 天可用率 · 99.94 %今天

    API 响应中位时间 240 ms

  • 短信下发
    正常
    90 天前90 天可用率 · 99.91 %今天
  • Webhook 投递
    正常
    90 天前90 天可用率 · 99.96 %今天

    API 响应中位时间 96 ms

所有时间均为 UTC。可用率按组件分别统计,并以受影响流量占比加权。

故障历史

维护窗口至少提前 48 小时在此公告。

请客服把你加入

2026年7月

性能下降已解决

Webhook deliveries failing to endpoints behind one certificate chain

受影响范围
Webhook 投递
持续时间
08:52–11:18 UTC · 2 小时 26 分钟
  1. 08:52 UTC

    Webhook deliveries to a subset of endpoints are failing with a TLS verification error. Events are queued and will be retried, so nothing is being dropped. Investigating.

  2. 09:14 UTC

    A base image update on the delivery workers shipped a trust store that no longer carries a cross-signed intermediate that several customer endpoints still serve. Around 4 % of registered endpoints are affected.

  3. 09:48 UTC

    We have pinned the previous trust store on the delivery workers and the queued deliveries are going out. The oldest queued event is 56 minutes old.

  4. 11:18 UTC

    Resolved. The queue is empty and every event from the window was delivered well inside the 24-hour retry budget. The trust store is now pinned explicitly and updated deliberately instead of being inherited from the base image.

部分中断已解决

Balance credits delayed by a payment callback backlog

受影响范围
支付(OxaPay)
持续时间
11:06–12:24 UTC · 1 小时 18 分钟
  1. 11:06 UTC

    Top-up credits are arriving late. Payments are being received and recorded, and no funds are at risk, but balances are updating with a delay. Investigating.

  2. 11:21 UTC

    OxaPay moved their callback traffic to a new source range this morning. Our allowlist rejected it, so callbacks were retried instead of accepted. About 340 callbacks are pending.

  3. 11:44 UTC

    The new range is allowlisted and the retries are being accepted. The backlog is draining oldest first at roughly 90 callbacks a minute.

  4. 12:24 UTC

    Resolved. Every pending credit has been applied and the longest delay was 78 minutes. We have asked to be notified of source-range changes in advance, and we now alert on a callback rejection rate above 1 % rather than only on queue depth.

2026年6月

维护中已解决

Singapore maintenance window overran

受影响范围
REST API, 控制台
持续时间
18:00–18:47 UTC · 47 分钟
  1. 18:00 UTC

    Maintenance window open, as announced on 2026-06-05. The Singapore entry point is draining and requests are being served from Frankfurt for the duration.

  2. 18:12 UTC

    The kernel upgrade is finished but one node is not rejoining the load balancer. Singapore stays drained while we look at it. Frankfurt is serving all traffic and latency from Asia is elevated to roughly 300 ms.

  3. 18:31 UTC

    The node was presenting a certificate issued before the upgrade. It has been reissued and the node is back in rotation.

  4. 18:47 UTC

    Resolved. Singapore is serving again at normal latency. We announced a 20-minute window and the drain lasted 47 minutes, which we should have said here sooner than we did.

部分中断已解决

Inbound SMS dropped on an Indonesian operator

受影响范围
短信下发, Webhook 投递
持续时间
05:44–12:26 UTC · 6 小时 42 分钟
  1. 05:44 UTC

    Activations on Indonesian numbers have been expiring without a code at a much higher rate than normal since about 05:10. Purchases and the rest of the catalogue are unaffected.

  2. 06:30 UTC

    The operator changed the route for inbound international SMS overnight without notice. Messages reach their network and are not forwarded to us. The affected ranges are out of routing as of 06:24.

  3. 07:05 UTC

    Indonesia is now served entirely by the two remaining operators. Success rate is 88 % against 93 % normally, and stock is about a third lower. Webhook volume for Indonesia is down accordingly.

  4. 09:40 UTC

    The operator has acknowledged the route change and is reverting it. We will not return the ranges to routing until our own probes have delivered cleanly for two hours.

  5. 12:26 UTC

    Resolved. The ranges are back in routing and delivering normally. 1,946 activations expired without a code during the incident; every one of them was refunded in full, automatically, with no ticket required.

2026年3月

性能下降已解决

Rate limiter rejecting requests that were inside their budget

受影响范围
REST API
持续时间
07:29–09:11 UTC · 1 小时 42 分钟
  1. 07:29 UTC

    A number of API keys are receiving 429 responses while comfortably inside their per-minute budget. Investigating.

  2. 07:48 UTC

    A counter shard lost its expiry during a cache node replacement overnight, so counts from the previous hour were never cleared for keys hashed onto that shard. Roughly 6 % of keys are affected.

  3. 08:15 UTC

    The affected shard has been flushed and the keys on it are being served normally. We are checking the remaining shards for the same condition.

  4. 09:11 UTC

    Resolved. Counters now carry an absolute expiry written at creation instead of one set after the first increment, so a node replacement cannot leave a counter without one. No account was charged for a request that was rejected.

2026年1月

部分中断已解决

Top-up invoice creation failing upstream

受影响范围
支付(OxaPay)
持续时间
10:22–13:05 UTC · 2 小时 43 分钟
  1. 10:22 UTC

    Creating a top-up invoice is failing with an upstream error for most currencies. Invoices already created are unaffected and payments already sent are being credited normally.

  2. 10:40 UTC

    OxaPay has confirmed an incident on their invoice API. Purchases, activations, rentals and the rest of the API are not affected; only creating a new top-up is.

  3. 11:35 UTC

    Partial recovery upstream. USDT-TRC20 and TON invoices are being created again. BTC and ETH still fail intermittently.

  4. 12:41 UTC

    All currencies are creating invoices again. We are keeping this open while we watch the error rate.

  5. 13:05 UTC

    Resolved. 214 invoice attempts failed during the window and none of them took a payment. The top-up page now names the specific currency that is unavailable instead of failing generically.

2025年11月

部分中断已解决

Dashboard failed to load after a deploy

受影响范围
控制台
持续时间
15:38–16:19 UTC · 41 分钟
  1. 15:38 UTC

    The dashboard is failing to load for some visitors with a chunk loading error. The API is unaffected, so automations are still running normally.

  2. 15:47 UTC

    A deploy at 15:31 invalidated the asset manifest while open sessions were still holding the previous one. Anyone who had the dashboard open before the deploy is affected; a hard reload works around it.

  3. 16:02 UTC

    The previous asset bundle has been re-published alongside the new one so both manifests resolve. Error reports have stopped.

  4. 16:19 UTC

    Resolved. Old asset bundles are now retained for 24 hours after a deploy, and the client reloads itself once when it sees a manifest it does not recognise.

2025年10月

部分中断已解决

Upstream operator degraded in Brazil

受影响范围
短信下发
持续时间
09:15–17:48 UTC · 8 小时 33 分钟
  1. 09:15 UTC

    Success rates on Brazilian numbers have fallen from 91 % to about 44 % since 08:30. Purchases still succeed, but a large share of activations are expiring without a code.

  2. 09:52 UTC

    One operator range is being rejected by several services at once, which normally means the range has been flagged rather than that delivery is broken. Its pool score is set to zero so routing avoids it.

  3. 11:10 UTC

    Brazilian success rate is back to 86 % on the remaining operators. Stock for Brazil is roughly 40 % below normal while the flagged range is excluded.

  4. 14:30 UTC

    The operator confirms the range was flagged by a downstream aggregator and is migrating the block. We are keeping it out of routing until our own probes show it recovering.

  5. 17:48 UTC

    Resolved. Brazil is at 90 % success on a reduced pool. Every activation that expired during the incident was refunded automatically. The flagged range stays out of routing until the migration is finished.

2025年8月

严重中断已解决

Primary database failover

受影响范围
REST API, 控制台, 支付(OxaPay), Webhook 投递
持续时间
04:07–04:52 UTC · 45 分钟
  1. 04:07 UTC

    The API is returning 500 for most requests and the dashboard will not load. The primary database is not responding. Investigating at the highest priority.

  2. 04:14 UTC

    The primary lost its storage volume. Automatic failover did not trigger because the primary was still answering health checks on its keepalive port. We are promoting the standby manually.

  3. 04:26 UTC

    The standby is promoted and the API is answering again. Error rate is back to baseline. We are now verifying whether any committed purchase was lost in the promotion.

  4. 04:52 UTC

    Resolved. 26 seconds of writes were lost in the promotion, covering 11 purchases. All 11 were refunded in full and the accounts were emailed individually. Health checks now run a real query instead of probing the keepalive port, and failover is exercised weekly.

2025年6月

维护中已解决

Scheduled maintenance: PostgreSQL major version upgrade

受影响范围
REST API, 控制台, 支付(OxaPay), Webhook 投递
持续时间
02:00–02:41 UTC · 41 分钟
  1. 02:00 UTC

    Maintenance window open, as announced on 2025-06-10. Writes return 503 with a Retry-After header; reads and long-polls continue to be served from the replica.

  2. 02:18 UTC

    The upgrade is complete and the primary is accepting writes. We are comparing query plans against the pre-upgrade baseline before calling this done.

  3. 02:41 UTC

    Resolved. Total write unavailability was 18 minutes, inside the 45-minute window we announced. Activations that were open during the window were extended by 18 minutes so that no one lost part of an activation to the maintenance.

2025年4月

部分中断已解决

Balance credits delayed by a stalled payment consumer

受影响范围
支付(OxaPay)
持续时间
18:41–21:05 UTC · 2 小时 24 分钟
  1. 18:41 UTC

    Top-ups paid in the last half hour are not appearing on balances. Payments are being received and recorded; it is the credit step that is backed up. No funds are at risk.

  2. 19:02 UTC

    The consumer that applies OxaPay callbacks to balances stopped acknowledging messages after a deploy at 18:10, and the queue has grown to about 400 callbacks. We are rolling that deploy back.

  3. 19:35 UTC

    The rollback is live and the backlog is draining at roughly 60 callbacks a minute. The oldest pending credit is 51 minutes old.

  4. 20:20 UTC

    Backlog cleared. Every payment received during the incident has been credited, including four invoices that expired while their callback was waiting.

  5. 21:05 UTC

    Resolved. The deploy removed the acknowledgement path for a callback that arrives for an invoice already in a terminal state, so those callbacks were retried forever and blocked the queue behind them. They are now acknowledged and logged.

2025年3月

性能下降已解决

Elevated SMS delivery times on Indian ranges

受影响范围
短信下发
持续时间
06:52–10:34 UTC · 3 小时 42 分钟
  1. 06:52 UTC

    Delivery times for Indian numbers are running above 60 seconds against a normal median of 9. Numbers are still being issued and the refund window is unchanged. Investigating.

  2. 07:20 UTC

    One of our two Indian upstreams is queueing inbound SMS at its own gateway. India is being routed to the second provider while we wait on their side.

  3. 08:05 UTC

    The shift is complete and new Indian activations are delivering at a median of 14 seconds. Activations bought before 07:20 that are still open will either deliver or auto-refund at the end of their window.

  4. 09:40 UTC

    The upstream has drained its queue and reports a full disk on a gateway node as the cause. We are keeping traffic on the second provider until their median has been stable for an hour.

  5. 10:34 UTC

    Resolved. India is back on both providers at a median of 9.2 seconds. 812 activations expired without a code during the incident and were refunded automatically.

接收通知

通过邮件或 Webhook 获取状态变更。

订阅 RSS 源