Enterprise Data Spillover System
The Enterprise Data Spillover System in NitroStack protects LLM context windows, client memory buffers, and network transports from payload exhaustion. When database queries, CSV exports, or log aggregations produce payloads exceeding configured limits (default: 10KB), the spillover interceptor automatically offloads full datasets to high-performance local storage, generates structured preview samples, and provides MCP resource pointers for on-demand inspection.
Table of Contents
- Overview & Architecture
- How Data Spillover Works
- Applying the Interceptor
- The Spillover Envelope
- Server Storage Drivers
- Client and AI Model Interaction
- Best Practices
Overview & Architecture
The Context Exhaustion Problem
When AI models call enterprise tools (such as database query executors or analytics collectors), responses often contain hundreds or thousands of rows totaling hundreds of kilobytes:
┌────────────────────────────────────────────────────────┐
│ Unprotected Tool Response Dump │
├────────────────────────────────────────────────────────┤
│ │
│ Model invokes: execute_sql("SELECT * FROM users") │
│ Response: 500KB JSON payload (125,000+ tokens) │
│ │
│ Consequences: │
│ • Model context window immediately filled │
│ • Token budget exhausted on a single invocation │
│ • Client buffer overflow and JSON parsing slowdowns │
│ • Model reasoning severely degraded by raw text data │
│ │
└────────────────────────────────────────────────────────┘
The NitroStack Solution
The DataSpilloverInterceptor sits directly on the tool execution pipeline, inspecting response byte lengths:
┌────────────────────────────────────────────────────────────────────────┐
│ Data Spillover Flow Architecture │
├────────────────────────────────────────────────────────────────────────┤
│ │
│ Tool Execution Result (e.g. 250KB JSON) │
│ │ │
│ ▼ │
│ DataSpilloverInterceptor │
│ │ │
│ ├── If payload size <= maxPayloadBytes: │
│ │ Return raw payload directly to LLM │
│ │ │
│ └── If payload size > maxPayloadBytes: │
│ 1. Persist full data into SpilloverStore (Memory or FS) │
│ 2. Generate truncated preview (first N items, record count│
│ 3. Return Spillover Envelope to LLM │
│ │
│ Returned Envelope: │
│ { │
│ "_spillover": true, │
│ "resourceUri": "resource://data-spillover/c837b92f-...", │
│ "preview": [ ... preview records ... ], │
│ "totalItems": 1500, │
│ "sizeBytes": 256000 │
│ } │
│ │
└────────────────────────────────────────────────────────────────────────┘
How Data Spillover Works
- Size Detection: When a tool handler completes, the interceptor calculates the serialized UTF-8 byte length. If within the threshold (
maxPayloadBytes, default: 10KB), the result passes through untouched. - Persistence: If the payload exceeds the threshold, the full dataset is written to the active
SpilloverStorewith a unique UUID key and a TTL expiration timestamp. - Smart Preview Generation:
generatePreview()inspects the data structure:- Arrays: Returns the first
Nelements (default: 3) with truncated fields and total item count. - Strings: Returns the first
Ncharacters (default: 500) with a[truncated]marker. - Objects: Extracts top-level keys with truncated values.
- Arrays: Returns the first
- Envelope Packaging: Returns a standardized
SpilloverEnvelopecontaining the preview, total record count, byte size, human summary, and an MCPresourceUri. - Resource Retrieval: The NitroStack server exposes a built-in MCP resource handler at
resource://data-spillover/{id}, enabling the client or model to read the complete dataset on demand.
Applying the Interceptor
The DataSpilloverInterceptor is applied via the @UseInterceptors() decorator. You can instantiate it directly with custom configuration:
import {
Controller,
ToolDecorator as Tool,
UseInterceptors,
ExecutionContext,
DataSpilloverInterceptor,
z,
} from '@nitrostack/core';
import { DatabaseService } from './database.service.js';
@Controller('database')
export class DatabaseController {
constructor(private readonly db: DatabaseService) {}
@Tool({
name: 'query_records',
description: 'Execute analytical queries with automatic context protection',
inputSchema: z.object({
sql: z.string().describe('SQL query string'),
}),
})
@UseInterceptors(
new DataSpilloverInterceptor({
maxPayloadBytes: 15 * 1024, // Spill over if payload > 15KB
previewItems: 3, // Include 3 sample rows in preview
previewStringChars: 300, // Truncate preview strings to 300 chars
spilloverTtlSeconds: 3600, // Keep resource cached for 1 hour
resourceUriPrefix: 'resource://data-spillover/',
})
)
async queryRecords(input: { sql: string }, ctx: ExecutionContext) {
ctx.logger.info('Executing analytical query', { sql: input.sql });
return await this.db.query(input.sql);
}
}
Static Configuration Helper
Alternatively, use the static factory helper:
@Tool({ name: 'export_csv' })
@UseInterceptors(
DataSpilloverInterceptor.configure({
maxPayloadBytes: 20 * 1024,
previewItems: 5,
})
)
async exportCsv(input: ExportInput, ctx: ExecutionContext) {
return await this.reports.generateCsv(input);
}
Interceptor Options Reference
| Option | Type | Default | Description |
|---|---|---|---|
maxPayloadBytes | number | 10240 (10KB) | Byte threshold before spillover triggers. |
spilloverTtlSeconds | number | 3600 (1 hour) | Time-to-live before spilled resources are pruned from storage. |
previewItems | number | 3 | Number of items included in array/object preview slices. |
previewStringChars | number | 500 | Maximum character length for preview strings before truncation. |
resourceUriPrefix | string | 'resource://data-spillover/' | URI scheme and prefix for the generated MCP resource. |
storage | 'memory' | 'filesystem' | SpilloverStore | 'memory' | Storage driver backend or custom implementation. |
storagePath | string | './.spillover_cache' | Directory path when using filesystem storage. |
The Spillover Envelope
When spillover triggers, the AI model receives a structured JSON object matching the SpilloverEnvelope schema:
{
"_spillover": true,
"resourceUri": "resource://data-spillover/b74dfa6e-4f1e-45fa-8025-a7b3e18c6109",
"mimeType": "application/json",
"sizeBytes": 284500,
"totalItems": 1250,
"summary": "Query returned 1,250 items (277.8 KB)",
"preview": [
{
"id": "usr_001",
"name": "Alice Johnson",
"email": "alice@example.com",
"status": "active"
},
{
"id": "usr_002",
"name": "Bob Smith",
"email": "bob@example.com",
"status": "pending"
},
{
"id": "usr_003",
"name": "Charlie Brown",
"email": "charlie@example.com",
"status": "active"
}
],
"hint": "The response exceeded the size limit and was saved as an MCP resource. Use resources/read with the resourceUri above to fetch the full data, or inspect the preview above."
}
Server Storage Drivers
NitroStack provides two built-in storage implementations:
1. Filesystem Storage (FsSpilloverStore)
Recommended for production workloads. Spools large datasets to disk, freeing Node.js heap memory:
import { createServer } from '@nitrostack/core';
const server = createServer({
name: 'enterprise-data-service',
version: '1.0.0',
spillover: {
driver: 'filesystem',
storageDir: './.spillover_cache',
maxSizeBytes: 200 * 1024 * 1024, // 200MB directory quota
},
});
- Atomic Disk Writes: Writes payloads to temporary files and renames them atomically to prevent partial reads.
- Auto-Pruning: Evicts expired records on background cleanup cycles or when approaching the storage byte cap.
- Heap Isolation: Large JSON strings are streamed directly to disk, avoiding V8 heap fragmentation.
2. In-Memory Storage (MemorySpilloverStore)
Ideal for development, lightweight microservices, and automated testing:
const devServer = createServer({
name: 'dev-service',
version: '1.0.0',
spillover: {
driver: 'memory',
maxSizeBytes: 25 * 1024 * 1024, // 25MB maximum in-memory storage
},
});
- Zero filesystem dependencies.
- High-speed in-memory Map lookup.
- Prunes expired records automatically on access and periodic sweeps.
Client and AI Model Interaction
When an AI model receives a _spillover: true envelope, it can handle the data intelligently:
- Immediate Analysis via Preview: The model uses the
preview,summary, andtotalItemsfields to answer high-level questions (e.g., verifying schema columns, validating status counts, or summarizing initial records). - Selective Full Retrieval: If the model requires complete data for downstream calculations, it calls the MCP standard
resources/readtool using the returnedresourceUri:
AI Model Request:
{
"method": "resources/read",
"params": {
"uri": "resource://data-spillover/b74dfa6e-4f1e-45fa-8025-a7b3e18c6109"
}
}
Server Response:
{
"contents": [
{
"uri": "resource://data-spillover/b74dfa6e-4f1e-45fa-8025-a7b3e18c6109",
"mimeType": "application/json",
"text": "[ { \"id\": \"usr_001\", ... complete dataset ... } ]"
}
]
}
Best Practices
- Set Realistic Thresholds: Set
maxPayloadBytesbased on typical model context limits. A threshold between 10KB and 25KB works best for enterprise environments. - Use Filesystem Storage in Production: For production servers serving multiple concurrent sessions, configure
filesystemstorage to preserve Node.js process heap memory. - Configure Appropriate TTLs: Align
spilloverTtlSecondswith session duration (e.g., 3600 seconds for 1-hour interactive sessions). - Combine with Code Mode: In high-throughput environments, combine
DataSpilloverInterceptorwithCodeModeTransformso models can process large datasets within local WASM scripts without dumping raw rows into context.