Stop feeding Claude
noisy context dumps.
ContextRail AI is developer infrastructure that routes semantic memory, compresses RAG payloads by up to 78%, and aligns Anthropic prompt caches — in under 50ms.
A high-velocity context highway
for autonomous agents
Standard RAG brute-forces entire document chunks into Claude's context. ContextRail intercepts the payload, applies salience reranking, retrieves cross-session episodic memory, and aligns Anthropic prompt cache boundaries — in <50ms.
Sub-50ms RAG Compression
Prunes syntactic fluff, redundant markdown, and low-salience paragraphs while preserving crucial technical needles. Up to 78% reduction, zero hallucination penalty.
Cross-Session Agent Memory
Equip Claude with persistent episodic memory. Our hybrid vector-and-knowledge-graph index remembers user decisions, codebase invariants, and tool results across sessions.
Prompt Cache Alignment
Formats prompts so static system guidelines sit cleanly behind Anthropic's 5-minute Prompt Caching breakpoints — up to 90% savings on repeat input tokens.
Zero Foundation Model Training
Your codebase and agent interactions are never used for training. Context is processed in volatile memory with AES-256 tenant isolation and GDPR sovereignty.
Native MCP Server
First-class compliance with Anthropic's open MCP standard. Drop our official server card into Claude Desktop or Claude Code in seconds.
Real-Time Token Telemetry
Inspect per-turn token economy, cache hit ratios, and millisecond routing latency from our developer console or via webhook alerts.
on multi-doc RAG threads
on European edge infra
on Claude Sonnet 4.6
per 10k agent turns
Connect to Claude Code & Desktop
in 60 seconds
ContextRail AI provides a verified Model Context Protocol server exposing real-time context routing and episodic memory tools directly to your Claude runtimes.
Expose ContextRail directly in Claude
By adding ContextRail AI to your claude_desktop_config.json, Claude gains automatic access to semantic memory queries and dynamic context pruning tools.
route_context — salience reranking & pruning
compress_rag_context — sub-50ms token reducer
retrieve_agent_memory — cross-session vector lookup
store_agent_memory — persistent episodic entity graph
$ claude mcp add --transport http contextrail https://api.contextrail.cloud/mcp
$ curl https://api.contextrail.cloud/.well-known/mcp/server-card.json
{
"mcpServers": {
"contextrail": {
"command": "npx",
"args": [
"-y",
"@contextrail/mcp-server"
],
"env": {
"CONTEXTRAIL_API_KEY":
"cr_live_your_key_here",
"CONTEXTRAIL_ENDPOINT":
"https://api.contextrail.cloud/mcp",
"TARGET_MODEL":
"claude-sonnet"
}
}
}
}
Measured performance on Claude Sonnet
Evaluated on real-world multi-repo refactoring and 100k+ token financial multi-document Q&A suites. ContextRail outperforms vanilla RAG and raw ingestion across all core dimensions.
| Routing Strategy | Context Reduction | Needle Recall | TTFT Latency | Cost / 10k turns |
|---|---|---|---|---|
|
ContextRail AI (MCP Engine) Semantic Reroute + Cache Breakpoint |
78.2% compressed | 99.4% accurate | 142 ms | $4.20 USD |
|
Standard Vector RAG (LangChain / LlamaIndex) Top-K chunk concatenation |
32.0% compressed | 76.5% accurate | 680 ms | $18.50 USD |
|
Raw Uncompressed Ingestion Full repo dump into 200k window |
0% (full dump) | 84.1% accurate | 1,850 ms | $38.90 USD |
Trusted by engineers shipping
Claude-powered products
Public HTTP API for programmatic
context routing
Every MCP tool is also a documented REST endpoint. Call ContextRail AI from any runtime, CI job, or backend service. TLS 1.3, bearer-key auth, RFC 9457 error bodies.
| Method | Path | Description | Auth | Rate Limit |
|---|---|---|---|---|
| POST | /api/v1/context/route |
Rank and prune a context payload before injection | Bearer key | 120 req/min |
| POST | /api/v1/context/compress |
Sub-50ms semantic RAG chunk compression | Bearer key | 120 req/min |
| POST | /api/v1/memory/query |
Cross-session episodic memory lookup | Bearer key | 60 req/min |
| POST | /api/v1/memory/write |
Persist agent decisions into the vector-graph index | Bearer key | 60 req/min |
| GET | /api/v1/metrics/efficiency |
Token savings, cache hit ratio, routing latency | Bearer key | 120 req/min |
| POST | /api/v1/webhooks |
Register an HTTPS sink for budget and latency alerts | Bearer key | 10 req/min |
Idempotency-Key support on all POST routes. MCP transport: http streamable at https://api.contextrail.cloud/mcp. Scale plan raises limits to 1,000 req/min.
Predictable plans for builders
and enterprise fleets
No credit tokens. Transparent monthly billing in USD with instant API key issuance and zero commitment.
Sandbox
For individual developers prototyping agents with Claude Code or local desktop scripts.
- 10,000 routed tokens/month
- 1 concurrent MCP agent
- Standard semantic compression
- Community Discord & GitHub support
Builder
For production startups building SaaS products on top of Claude Sonnet & Haiku.
- 2,500,000 routed tokens/month
- Up to 10 active MCP agents
- Semantic vector agent memory
- Anthropic Prompt Cache alignment
- Email support (sub-24h)
Scale
For heavy multi-agent fleets, high-frequency autonomous workflows, and enterprise scale.
- 15,000,000 routed tokens/month
- Unlimited concurrent MCP agent fleets
- Hybrid vector + knowledge graph memory
- Dedicated priority routing & webhooks
- 99.9% Uptime SLA & Slack eng. channel
Technical specs & FAQs
Everything about ContextRail integration, security guarantees, and Claude model compatibility.
@contextrail/mcp-server) and exposes a verified server card at /.well-known/mcp/server-card.json. Paste the snippet into your Claude configuration file and your assistant immediately acquires our routing and memory tools.
Building high-throughput infrastructure
for the age of autonomous agents
Why we built this
ContextRail AI Software S.L. is a pure B2B software product engineered and operated by our systems team at Paseo de la Castellana 95, 28046 Madrid, Spain. We believe the biggest bottleneck facing autonomous agents is not model intelligence, but context bandwidth and retrieval bloat.
By treating context as a high-speed routing fabric rather than a static text bucket, we empower engineering teams to run complex multi-agent architectures on Claude with predictable costs, instant retrieval, and cross-session persistence.
David Morales Santillana
Ready to stop burning tokens
on noise?
Deploy your Sandbox environment in 90 seconds. No credit card. No lock-in. Start routing smarter context to Claude today.