System status
Real-time status of Predicue services and infrastructure. We strive for 99.9% uptime for all services.
Service components
Web panel
Working - 99.98% uptime over the last 90 days. All pages load normally.
REST API
Working - 99.95% uptime over the last 90 days. Average response time: 142 ms.
Market scanner
Live - Real-time data feed active, processing 400+ markets. Data relevance: <30 s.
Sending letters
Working - Bulletins and notifications are delivered normally. Average delivery time: 2.3 seconds.
Polymarket Data Channel
Working - Connected and syncing. Last sync: just now.
Webhooks
Working - All webhook destinations receive events normally. Average delivery: 1.8 seconds.
Recent Incidents
June 18, 2026 - Delay in delivery of reports (Resolved)
Duration: 45 minutes · Influence: Delay in sending email summaries for ~15% of users
Cause: Limit SendGrid request rate during peak delivery window (8:00 AM EST). Our burst of emails exceeded their minute limit.
Solution: A backup email provider (Postmark) has been implemented. Delivery now automatically switches when the main provider's limit approaches.
Prevention: Distributing bulletins over a 15-minute window instead of sending them all at once.
June 3, 2026 - API Latency Surge (Solved)
Duration: 22 minutes · Influence: API response time increased to 2-3 seconds
Cause: Exhaustion of the database connection pool during periods of high traffic. A surge of simultaneous requests occupied all available connections.
Solution: The connection pool has been increased from 50 to 200, and reuse of connections with a maximum lifetime of 5 minutes has been added.
Prevention: Added monitoring notifications when the pool load is 70%.
May 21, 2026 - Scanner data obsolescence (Solved)
Duration: 12 minutes · Influence: The market scanner showed data with a delay of up to 5 minutes
Cause: Intermittent Polymarket API timeouts causing the data pipeline to lag.
Solution: Implemented Circuit Breaker pattern for Polymarket API calls with automatic transition to cached data.
May 8, 2026 - Summary generation failure (Solved)
Duration: 3 hours · Influence: Morning bulletins were not generated for ~30% of users
Cause: Increased AI model inference timeout due to larger context window. Tasks were canceled before completion.
Solution: Increased inference timeout and optimized context window size. Added retry logic for failed tasks.
Uptime obligations
We are targeting 99.9% uptime for all services, measured monthly. This allows for approximately 43 minutes of planned or unplanned downtime per month.
SLA for Life plan
Life plan holders receive SLA compensation for any downtime exceeding 0.1% per calendar month:
- 99.9%+ uptime: No compensation
- 99.0% - 99.9%: 10% credit toward account balance
- 95.0% - 99.0%: 25% credit
- Below 95.0%: 50% credit + priority incident analysis
Maintenance windows
Scheduled maintenance occurs on Sundays between 2:00-4:00 AM EST. We notify users at least 48 hours in advance via email and in-app. During maintenance, the panel may be unavailable for a short period of time (typically <5 minutes).
Status Updates
During any incident, status updates are posted here and on our X (Twitter) account (@PredicueAI) every 15 minutes until fully resolved. Incident reports are published within 48 hours.
How we monitor
We use a multi-layered monitoring approach to quickly detect and respond to problems:
- Infrastructure monitoring: Datadog for server metrics, network performance and resource utilization
- Application monitoring: Sentry for bug and performance tracking
- Uptime monitoring: Pingdom for external accessibility checks from 10+ locations around the world
- Data pipeline monitoring: Custom alerts for freshness, completeness and accuracy of data
- Duties: The engineering team is on duty 24/7 with a critical incident response SLA of 15 minutes