How to Build a Real‑Time Telemetry Pipeline for SpaceX Starship Launches
September 28, 2026· 9 min read
TL;DR: A robust, cloud‑native streaming pipeline can ingest, process, and monetize the bursty telemetry from SpaceX’s Starship launch on Sep 28 2026 while keeping AI‑driven safety checks deterministic and cost‑controlled.
- Introduction
SpaceX’s Starship is poised to become the world’s first fully reusable orbital launch vehicle that also serves as a revenue‑generating platform. The upcoming Sep 28 2026 flight will be the first launch where SpaceX intends to sell high‑resolution telemetry to downstream customers and license its AI‑based safety analytics as a service.
From a data‑engineering perspective, the event is a stress test: a 30‑minute burst that can exceed 5 Gbps of raw downlink, translating to ≈ 2 TB of telemetry in a single pass. The data must be:
✔️Ingested without loss – packet‑level reliability is mandatory because each sensor reading can be the difference between a successful ascent and an abort.
✔️Processed in sub‑second latency – AI safety models must flag anomalies within ≤ 50 ms to influence abort decisions before the vehicle reaches 100 km altitude.
✔️Monetized in real time – customers expect usage‑based billing that reflects exactly what they consume, not a post‑hoc estimate.
This guide walks you through a production‑grade architecture that satisfies those constraints, explains the trade‑offs of major technology choices, and provides concrete implementation details you can copy‑paste into your own environment.
Small packets increase per‑record overhead; efficient serialization is crucial.
Sequence number
64‑bit monotonically increasing
Enables deduplication and replay detection.
Timestamp
UTC, nanosecond precision
Needed for deterministic ordering across shards.
Peak throughput
5 Gbps (≈ 1.9 TB/min)
Drives shard count, network sizing, and autoscaling policies.
Burst duration
30 minutes (ascent + coast)
Determines total data volume (~2 TB) and storage tiering strategy.
Error rate
Up to 2 % packet loss due to RF fading
Must be compensated by retransmission handling and idempotent downstream logic.
These characteristics dictate the design of every layer in the pipeline—from the edge gateway that talks to the RF antenna to the downstream analytics that feed the abort decision.
- Architectural Overview
+-------------------+ +---------------------+ +-------------------+
| RF Antenna Array | ---> | Protocol‑Translation| ---> | Cloud‑Native |
| (S‑/X‑band) | | Gateway (Rust) | | Streaming Service|
| |
| (JSON / Protobuf) |
v v
+-------------------+ +-------------------+
| Ingestion Layer | ---> | Real‑Time AI |
| (Kinesis / Kafka)| | Inference (Triton)|
| Enriched Telemetry |
| Enrichment Store | ---> | Billing & Usage |
| (ClickHouse) | | API (OpenAPI) |
+-------------------+
| Long‑Term Archive |
| (Glacier Deep) |
✔️Protocol‑Translation Gateway – a low‑latency Rust service that decodes CCSDS frames, adds schema metadata, and writes JSON/Protobuf records to the streaming service.
✔️Ingestion Layer – either Amazon Kinesis Data Streams v2 (managed, serverless) or self‑managed Apache Kafka on Amazon EKS/EKS‑managed node groups. Both support horizontal scaling, but Kafka gives finer‑grained control over partition placement and retention.
✔️Real‑Time AI Inference – NVIDIA Triton Inference Server (or TorchServe) running TensorRT‑optimized ONNX models, consuming the same consumer group as downstream analytics to guarantee exactly‑once scoring.
✔️Enrichment Store – ClickHouse for fast columnar queries on annotated telemetry; also serves as the source for the per‑customer usage tables.
✔️Billing & Usage API – a lightweight Go/Node service exposing OpenAPI 3.0 endpoints that write usage rows to ClickHouse in near‑real time.
✔️Long‑Term Archive – AWS Glacier Deep Archive for raw payload after a 24‑hour hot‑store window.
The following sections dive into each component, provide concrete configuration snippets, and discuss the trade‑offs you’ll encounter.
- Designing a Scalable Ingestion Layer
4.1 Choosing Between Kinesis and Kafka
Feature
Amazon Kinesis v2
Apache Kafka on EKS
Managed vs Self‑Managed
Fully managed, no cluster ops
Requires ops (EKS, Helm)
Shard/Partition Scaling
Autoscaling via OnDemand or Enhanced Fan‑Out; max 10 000 shards per stream
Kafka Autoscaler (KAS) can add partitions on the fly; limited by broker count
Throughput Guarantees
1 MB/s per shard (≈ 8 Mbps)
Depends on broker hardware; typical 10 Gbps per broker with SSDs
EC2/EKS instance + EBS cost; cheaper at high sustained throughput
Latency
~30 ms median (depends on region)
Sub‑10 ms intra‑AZ, higher cross‑AZ
Recommendation: For a single‑launch, burst‑only workload, Kinesis offers the fastest path to production because you avoid cluster management. However, if you anticipate continuous high‑throughput streams (e.g., multiple launches per week) or need fine‑grained retention policies, Kafka on EKS gives you more flexibility and lower per‑GB cost.
4.2 Provisioning the Required Capacity
Kinesis Example (5 Gbps ≈ 625 MB/s):
bash
aws kinesis create-stream \
--stream-name starship-telemetry \
--shard-count 750 \
--stream-mode ON_DEMAND
Kafka Example (3 brokers, 2 TB total storage):
yaml
replicaCount: 3
resources:
limits:
cpu: "8"
memory: "32Gi"
requests:
cpu: "4"
memory: "16Gi"
storage:
type: gp3
size: 4Ti # 4 TB per broker → 12 TB total (room for replication)
The Kafka Autoscaler (KAS) can be configured to add partitions when the producer lag exceeds a threshold:
yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: kafka-partition-scaler
spec:
scaleTargetRef:
name: kafka-broker
triggers:
- type: kafka
bootstrapServers: kafka:9092
topic: starship-telemetry
lagThreshold: "5000000" # 5 M messages
activationLagThreshold: "1000000"
4.3 Back‑Pressure and Local Buffering
Even with autoscaling, the first few seconds of a launch can outpace provisioning. The gateway must therefore hold data locally:
✔️NVMe Cache – 200 GB on the edge node can buffer ~30 seconds at 5 Gbps.
✔️Ring Buffer Implementation – a lock‑free circular buffer (e.g., crossbeam::queue::ArrayQueue in Rust) provides O(1) enqueue/dequeue with minimal CPU overhead.
rust
js
let buffer = ArrayQueue::<TelemetryRecord>::new(2_000_000); // 2M records ≈ 200 GB
loop {
match decode_ccsds_frame(&mut socket) {
Ok(rec) => {
if buffer.push(rec).is_err() {
// Buffer full → trigger autoscaler via HTTP call
trigger_autoscale().await;
}
}
Err(e) => log::warn!("Decode error: {}", e),
}
}
When the buffer reaches 80 % occupancy, the service should emit a CloudWatch metric (GatewayBufferUtilization) that the autoscaler watches. This prevents uncontrolled spill‑over and gives you a deterministic safety valve.
- Protocol‑Translation Gateway
SpaceX’s downlink uses CCSDS Space Packet Protocol with custom framing. The gateway’s responsibilities:
Schema enrichment – attach a Protobuf schema ID (telemetry.v1) and a schema version field.
Timestamp normalization – convert spacecraft‑local time to UTC nanoseconds using the embedded GPS week number.
Deduplication token – compute a hash of sequencenumber || payload (e.g., SHA‑256 truncated to 64 bits) and embed it as dedupid.
Why Rust?
✔️Zero‑cost abstractions keep CPU usage < 10 % on an m6i.large (2 vCPU, 8 GiB) while handling > 10 M packets/s.
✔️Predictable memory layout eliminates GC s that could add jitter to the latency budget.
5.1 Publishing to the Stream
Both Kinesis and Kafka have high‑throughput producer libraries. The gateway should use batching (e.g., 500 KB per request) and asynchronous I/O to keep the network pipe full.
rust
js
let producer = KinesisProducer::new("starship-telemetry");
let batch = telemetry_records.iter()
.map(|r| r.to_json())
.collect::<Vec<_>>();
producer.put_records(batch).await?;
For Kafka, enable linger.ms = 5 and batch.size = 1 MiB to balance latency vs throughput.
- Idempotency and Deduplication
Duplicate packets are inevitable because the RF link may retransmit lost frames. A single‑pass deduplication strategy avoids costly downstream re‑processing:
Hash‑based key – dedup_id (64 bits) becomes the primary key in the downstream ClickHouse table.
Exactly‑once consumer – use Kafka’s transactional consumer (isolation.level=read_committed) or Kinesis’s deduplication token (ExplicitHashKey).
Side‑effect‑free processing – all enrichment steps (e.g., model inference) must be pure functions of the input record; otherwise, a duplicate could cause double‑counted billing.
ClickHouse Table Example:
sql
CREATE TABLE telemetry_raw (
dedup_id UInt64,
seq_num UInt64,
ts DateTime64(9, 'UTC'),
payload JSON,
PRIMARY KEY dedup_id
) ENGINE = MergeTree()
ORDER BY dedup_id;
If an insert arrives with an existing dedup_id, ClickHouse will ignore it (using INSERT ... ON CONFLICT DO NOTHING semantics via the ReplacingMergeTree engine).
Disable dynamic shapes, use FP16 for consistency, set CUDALAUNCHBLOCKING=1 in the Triton container.
Version control
Store models in an S3 bucket with semantic versioning (model/v1.2.3/). Triton loads from a manifest file that can be atomically swapped.
Explainability
Enable SHAP integration in Triton via a custom Python backend; emit shap_values to a side‑channel topic.
7.2 Deploying the Inference Service
Dockerfile (Triton + Python backend):
dockerfile
FROM nvcr.io/nvidia/tritonserver:23.09-py3
COPY model /models/starship_anomaly/
COPY shap_backend.py /opt/tritonserver/backends/shap/
ENV TRITON_MODEL_REPOSITORY=/models
ENV SHAP_MODEL_PATH=/models/shap/
Kubernetes Deployment (GPU‑enabled):
yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: triton-inference
spec:
replicas: 2
selector:
matchLabels:
app: triton
template:
metadata:
labels:
app: triton
spec:
containers:
- name: triton
image: myrepo/triton:latest
resources:
limits:
nvidia.com/gpu: "1"
env:
- name: TRITON_MODEL_CONTROL_MODE
value: "explicit"
- name: TRITON_MODEL_REPOSITORY
value: "/models"
7.3 Explainability Side‑Channel
The SHAP values are written to a dedicated Kafka topicstarship-shap. Downstream auditors can replay this topic to reconstruct why a particular anomaly flag was raised.
json
{
"dedup_id": "0x1A2B3C4D",
"prediction": "ANOMALY",
"shap": {
"thrust": 0.42,
"temperature": -0.15,
"vibration": 0.23,
"timestamp": "2026-09-28T12:34:56.789123456Z"
}
}
Storing these values in ClickHouse enables SQL‑based audit queries:
sql
SELECT dedup_id, prediction, shap['thrust'] AS thrust_shap
FROM shap_events
WHERE prediction = 'ANOMALY'
ORDER BY timestamp DESC
LIMIT 10;
- Enrichment, Storage, and Query Layer
8.1 ClickHouse for Fast Enriched Queries
ClickHouse’s columnar storage and vector‑index capabilities make it ideal for:
✔️Time‑series analytics – e.g., “average thrust per second during max‑Q”.
✔️Ad‑hoc anomaly investigations – low‑latency joins between raw telemetry and SHAP values.
✔️Billing aggregation – per‑customer usage can be summed in milliseconds.
Schema Sketch:
sql
CREATE TABLE telemetry_enriched (
customer_id UInt32,
thrust Float32,
temperature Float32,
vibration Float32,
anomaly UInt8, -- 0 = OK, 1 = flagged
shap JSON,
PRIMARY KEY (customer_id, ts)
) ENGINE = MergeTree()
ORDER BY (customer_id, ts);
8.2 Data Retention Policy
Tier
Duration
Storage Class
Cost (2026 US‑East‑1)
Hot
24 h
S3 Standard
$0.023/GB‑mo
Warm
30 d
S3 Intelligent‑Tiering
$0.0125/GB‑mo
Cold
1 y
Glacier Flexible Retrieval
$0.004/GB‑mo
Deep Archive
1 y
Glacier Deep Archive
$0.00099/GB‑mo
A Lifecycle Rule on the S3 bucket automatically transitions objects after the defined periods, ensuring that the 2 TB raw payload costs less than $2 k after the first day.
- Monetizing Launch Telemetry
9.1 Pricing Model
Product
Unit
Price
Example Revenue (2 TB launch)
Raw Telemetry
MB
$0.12
2 TB = 2 048 000 MB → $245 760
Enriched Telemetry
MB
$0.35
30 % uptake → 614 400 MB → $215 040
Inference Calls
per call
$0.025
10 M calls → $250 000
Analytics Dashboard
per seat/month
$1 500
5 seats → $7 500
Total potential revenue per launch: ≈ $718 k (assuming 30 % enriched uptake).
9.2 Billing‑Ready API
Expose a RESTful OpenAPI 3.0 service that:
✔️Accepts OAuth2 bearer tokens per customer.
✔️Provides /usage/record endpoint that receives a JSON payload { "dedup_id": "...", "bytes": 1024 }.
✔️Writes directly to a high‑throughput ClickHouse tablecustomer_usage.
OpenAPI snippet:
yaml
paths:
/usage/record:
post:
security:
- oauth2: [usage:write]
requestBody:
required: true
content:
application/json:
schema:
$ref: '#/components/schemas/UsageRecord'
responses:
'202':
description: Accepted
components:
securitySchemes:
oauth2:
type: oauth2
flows:
clientCredentials:
tokenUrl: https://auth.example.com/oauth2/token
scopes:
usage:write: Record telemetry usage
schemas:
UsageRecord:
type: object
required: [dedup_id, bytes]
properties:
dedup_id:
type: string
format: uuid
bytes:
type: integer
minimum: 1
The API can be rate‑limited at 10 k requests per second (well below the 5 Gbps ingest) and autoscaled using AWS API Gateway + Lambda for the thin façade, while the heavy write path goes straight to ClickHouse via a gRPC bulk endpoint.
- Cost Management and Cloud Budgeting
10.1 Predicting the Spend
Component
Unit Cost (2026)
Expected Usage
Approx. Cost
Kinesis shards (750)
$0.015/shard‑hr
0.5 hr (burst)
$5.6
Kinesis PUT payload
$0.014 per GB
2 TB
$28
EMR Spark (spot, 70 % fleet)
$0.10 per vCPU‑hr
200 vCPU‑hrs
$20
Triton GPU (p4d.24xlarge)
$32 per hr
1 hr
$32
ClickHouse (c5.4xlarge)
$0.68 per hr
2 hr
$1.4
S3 Standard (24 h)
$0.023/GB‑mo
2 TB
$1.1
Glacier Deep Archive (30 d)
$0.00099/GB‑mo
2 TB
$0.06
Total
≈ $88
The burst cost is modest compared to the potential revenue. However, mis‑configured autoscaling can cause runaway shard creation. To safeguard:
✔️Set a hard shard‑count ceiling (maxShards = 800).
✔️Enable CloudWatch alarms on WriteProvisionedThroughputExceeded > 80 % for > 5 min → trigger a Lambda that s the stream (UpdateShardCount to 0) and notifies the ops team.
✔️Tag all resources with Project=StarshipTelemetry and Owner=LaunchTeam for cost allocation reports.
10.2 Spot‑Instance Trade‑offs
Spot instances provide up to 70 % discount but can be reclaimed with a 2‑minute warning. For real‑time inference, you cannot rely on spot; use on‑demand GPU for Triton. For batch enrichment (e.g., writing to ClickHouse, secondary analytics), spot is safe because the pipeline is idempotent and can resume after a brief interruption.
- Monitoring, Alerting, and Observability
Metric
Source
Alert Threshold
GatewayBufferUtilization
Prometheus (gateway)
80 % for > 30 s
KinesisWriteProvisionedThroughputExceeded
CloudWatch
90 % for > 1 min
InferenceLatencyMs
Triton metrics endpoint
45 ms for > 5 % of calls
FalsePositiveRate
Custom Prometheus rule (based on SHAP)
0.2 % for > 2 min
ClickHouseInsertLatency
ClickHouse metrics
100 ms for > 1 % of inserts
Grafana Dashboards should display:
✔️Ingress rate (Gbps) per shard/partition.
✔️End‑to‑end latency from antenna to enriched ClickHouse row.
✔️Model health – confusion matrix updates every 10 seconds.
All alerts feed into a PagerDuty service with two escalation paths: SRE on‑call for infrastructure alerts, ML Ops for model‑drift alerts.
- Testing and Validation
12.1 Load‑Testing the Ingestion Path
Generate synthetic CCSDS frames using a Python script that mimics the real packet distribution (payload size, sequence gaps).
Replay at 5 Gbps using tcpreplay on a dedicated EC2 instance (c5n.18xlarge).
Decouples heavy explainability compute from main path
Additional storage & latency for SHAP data
When auditability is a regulatory requirement
- Security and Compliance
Transport Encryption – Use TLS 1.3 for all connections (gateway → Kinesis/Kafka, Triton → downstream).
At‑Rest Encryption – Enable SSE‑KMS on S3 buckets and Transparent Data Encryption (TDE) on ClickHouse.
IAM Least‑Privilege – Grant the gateway only kinesis:PutRecord on the specific stream; Triton pods receive a service‑account with s3:GetObject for model assets.
Audit Logging – Forward CloudTrail events to a dedicated audit log in Elasticsearch; retain for 2 years to satisfy aerospace regulator requirements.
Data Residency – Keep all raw telemetry in the US‑East‑1 region to comply with SpaceX’s data‑location policy.
- Disaster Recovery and High Availability
Failure Mode
Recovery Strategy
Gateway node crash
Deploy two identical gateway pods behind an Elastic Load Balancer; use sticky sessions based on RF antenna ID.
Kinesis shard throttling
Autoscaler adds shards; alarm triggers fallback to on‑prem buffer (local SSD) and retries.
Kafka broker loss
Enable replication factor = 3; Zookeeper (or KRaft) automatically elects a new leader.
Triton GPU node failure
Run two replicas behind a ClusterIP Service; health checks route traffic away from the failing pod.
ClickHouse node outage
Use ReplicatedMergeTree with at least two replicas; queries automatically failover.
Regional AWS outage
Replicate the entire pipeline in US‑West‑2 using Cross‑Region Replication for the S3 bucket; failover via Route 53 latency‑based routing.
All failover actions should be drill‑tested at least quarterly.
- Operational Playbook for Launch Day
Time (UTC)
Action
Owner
T‑00:30
Verify gateway health (GatewayBufferUtilization < 30 %).
Edge Ops
T‑00:20
Warm‑up Triton GPU nodes (run a 5‑minute inference warm‑up batch).
ML Ops
T‑00:10
Enable Kinesis on‑demand scaling and set shard‑count ceiling.
Cloud FinOps
T‑00:05
Start billing API and confirm ClickHouse write latency < 50 ms.
Billing Team
T‑00:00
Launch begins – monitor IngressRate dashboard.
SRE Lead
T + 02 min
First anomaly flag expected – verify alert appears in Mission Control UI.
Safety Team
T + 10 min
Check spot‑instance pre‑emptions; if any, confirm Spark jobs resume.
Data Engineering
T + 30 min
Burst ends – automatically transition raw data to Glacier Deep Archive.
Storage Ops
T + 45 min
Run post‑flight billing reconciliation script; compare usage counters to expected 2 TB.
Finance
T + 60 min
Conduct post‑mortem meeting; capture SLA metrics and any incidents.
All Stakeholders
A single source of truth (a Confluence page with run‑books) ensures every team knows their exact responsibilities.
- Conclusion
The Sep 28 2026 Starship launch is more than a spectacular aerospace event; it is a real‑time data challenge that forces engineers to adopt streaming‑first architectures, deterministic AI inference, and usage‑based monetization from day one. By:
✔️Choosing a horizontally scalable ingestion service (Kinesis or Kafka) with automatic shard/partition scaling,
✔️Deploying a low‑latency Rust gateway that decodes CCSDS frames, handles back‑pressure, and guarantees idempotency,
✔️Embedding a TensorRT‑optimized inference engine directly in the ingestion path, with canary rollouts and SHAP explainability,
✔️Storing enriched telemetry in ClickHouse for sub‑second analytics and billing,
✔️Applying rigorous cost controls, monitoring, and disaster‑recovery practices,
you can not only survive the 5 Gbps, 2‑TB burst but also turn it into a $0.5 M+ revenue stream per launch.
Treat the telemetry as a product, not a by‑product. The architecture described here will serve you for Starship and for any future high‑stakes aerospace or industrial IoT scenario where milliseconds matter and data is money.
- References
✔️Forbes – “SpaceX’s Starship Rocket Is Heading To Orbit–With A Lot On The Line” (Sep 26 2026). External resource
✔️Yahoo Finance – “SpaceX’s Next Starship Launch Will Generate Revenue — Sort Of” (Sep 2026). External resource
✔️AWS Documentation – Kinesis Data Streams Scaling and Pricing. External resource
✔️NVIDIA Triton Inference Server – Deployment Guide (v2.41). External resource
What streaming service can handle the 5 Gbps telemetry burst from Starship?+
A cloud‑native service like Amazon Kinesis Data Streams v2 or a self‑managed Kafka cluster on EKS with the Kafka Autoscaler can scale to the required shard count; Kinesis needs ~5 000 shards, while Kafka can add partitions dynamically.
How do I keep AI inference latency under the 50 ms abort‑decision window?+
Deploy the model on NVIDIA Triton with TensorRT‑optimized ONNX, use a batch size of 1, pre‑allocate tensors, and colocate the inference service in the same consumer group as the ingestion pipeline to guarantee exactly‑once processing.
Can telemetry data from a Starship launch be monetized?+
Yes. SpaceX plans to charge $0.12/MB for raw telemetry and $0.35/MB for enriched streams, plus $0.025 per inference call for the safety analytics platform, potentially generating $0.5 M+ per launch.
What cost‑control measures should I implement for a one‑off launch event?+
Set shard‑count caps, use spot instances for downstream processing, archive raw data to Glacier Deep Archive after 24 hours, and embed Prometheus alerts on resource utilization.
Why is it important to treat telemetry as a product rather than a by‑product?+
Treating telemetry as a product forces you to build billing‑ready APIs, enforce data quality, and capture revenue streams; otherwise you risk retrofitting a batch pipeline after the launch and missing commercial opportunities.
The week's best on engineering, AI, and security — one email, no noise.
Read next
Related topicbusiness tech·September 22, 2026
How to Fix Massive Xbox Game Downloads with Efficient Asset Management
TL;DR: Use modular installation, aggressive compression, and streaming assets to keep Xbox game downloads under 100 GB, avoiding bandwidth bottlenecks and stora