Designing Scalable AI Copilots: Architecture and Performance
Introduction
AI copilots represent a distinct interface paradigm—neither standalone chat applications nor traditional command interfaces. A copilot augments the user's primary workflow, providing contextual assistance, suggestions, and actions without disrupting the main task. Think GitHub Copilot in VS Code, Notion AI inline, or Figma's AI assistant.
The architectural challenge is integration depth. A copilot must deeply understand the host application's state, respond to context changes in real-time, inject suggestions at appropriate moments, and execute actions within the application's domain—all while remaining unobtrusive. This requires tight coupling between the AI system and the host application's architecture in ways that chat interfaces don't demand.
GIF via GIPHY
Production copilot systems at scale face unique challenges: maintaining sub-100ms suggestion latency, managing context that spans documents and user history, handling the inherent conflict between proactive suggestions and user autonomy, and building trust through predictable, explainable behavior.
Scale Context
Modern AI copilot systems operate under demanding performance constraints:
| Metric | Production Scale |
|---|---|
| Active copilot sessions | 100K-1M |
| Suggestion requests/second | 10K-100K |
| P95 suggestion latency | <200ms |
| Context window per session | 8K-32K tokens |
| Suggestions per user/hour | 50-500 |
| Acceptance rate | 20-40% |
| Context updates/second | 10-100 per session |
| Concurrent inline completions | 1-5 per user |
| Undo/revision rate | 15-30% |
| Background pre-computation | 1-10 requests/second/user |
GIF via GIPHY
The critical requirement: suggestions must feel instantaneous while being contextually accurate. Users won't wait for AI—the copilot must predict and pre-compute.
High-Level Architecture
┌─────────────────────────────────────────────────────────────────────────────┐
│ Host Application │
│ ┌───────────────────────────────────────────────────────────────────────┐ │
│ │ Application UI Layer │ │
│ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │
│ │ │ Editor │ │ Canvas │ │ Form │ │ Table │ │ │
│ │ │Component │ │Component │ │Component │ │Component │ │ │
│ │ └────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘ │ │
│ │ │ │ │ │ │ │
│ │ ┌────┴─────────────┴─────────────┴─────────────┴────────────────┐ │ │
│ │ │ Copilot Integration Layer │ │ │
│ │ │ ┌───────────┐ ┌───────────┐ ┌───────────┐ ┌───────────┐ │ │ │
│ │ │ │ Context │ │Suggestion │ │ Action │ │ Ghost │ │ │ │
│ │ │ │ Collector │ │ Renderer │ │ Executor │ │ Text │ │ │ │
│ │ │ └───────────┘ └───────────┘ └───────────┘ └───────────┘ │ │ │
│ │ └───────────────────────────┬───────────────────────────────────┘ │ │
│ └──────────────────────────────┼───────────────────────────────────────┘ │
│ │ │
│ ┌──────────────────────────────┴───────────────────────────────────────┐ │
│ │ Copilot Core Engine │ │
│ │ ┌───────────┐ ┌───────────┐ ┌───────────┐ ┌───────────┐ │ │
│ │ │ Context │ │Prediction │ │ Cache │ │ Trigger │ │ │
│ │ │ Manager │ │ Engine │ │ Manager │ │ Engine │ │ │
│ │ └───────────┘ └───────────┘ └───────────┘ └───────────┘ │ │
│ └──────────────────────────────┬───────────────────────────────────────┘ │
└─────────────────────────────────┼───────────────────────────────────────────┘
│
┌──────────────┼──────────────┐
▼ ▼ ▼
┌───────────┐ ┌───────────┐ ┌───────────┐
│ Copilot │ │ Edge │ │ Context │
│ BFF │ │ Cache │ │ Service │
└─────┬─────┘ └───────────┘ └───────────┘
│
▼
┌─────────────────────────────────────────┐
│ AI Inference Layer │
│ ┌─────────┐ ┌─────────┐ ┌─────────┐ │
│ │ Fast │ │ Quality │ │Embedding│ │
│ │ Model │ │ Model │ │ Model │ │
│ └─────────┘ └─────────┘ └─────────┘ │
└─────────────────────────────────────────┘
Suggestion Lifecycle
Context Change → Trigger Evaluation → Context Assembly →
Speculative Fetch → Cache Check → Model Inference →
Suggestion Ranking → UI Rendering → User Decision →
Acceptance/Rejection → Learning Signal
GIF via GIPHY
The fundamental insight: copilots must predict user intent before explicit requests. The architecture optimizes for speculation and caching, not just response time.
Context Collection Architecture
Multi-Source Context Aggregation
interface CopilotContext {
// Immediate context - what user is doing right now
immediate: {
cursorPosition: Position;
selection: Selection | null;
currentLine: string;
surroundingLines: string[];
activeElement: ElementContext;
};
// Document context - broader document understanding
document: {
type: DocumentType;
structure: DocumentStructure;
recentChanges: Change[];
semanticSections: SemanticSection[];
};
// Project context - workspace-level understanding
project: {
fileStructure: FileTree;
dependencies: Dependency[];
conventions: Convention[];
relatedFiles: RelatedFile[];
};
// User context - personalization signals
user: {
recentActions: Action[];
preferences: Preference[];
acceptancePatterns: AcceptancePattern[];
skillLevel: SkillLevel;
};
// Session context - current work session
session: {
goals: InferredGoal[];
workingSet: WorkingSetFile[];
recentSuggestions: SuggestionHistory[];
};
}
class ContextCollector {
private collectors: Map<ContextType, ContextSourceCollector>;
private contextCache: LRUCache<string, ContextFragment>;
private updateDebouncer: Map<ContextType, ReturnType<typeof debounce>>;
constructor(config: CollectorConfig) {
this.collectors = new Map([
['immediate', new ImmediateContextCollector()],
['document', new DocumentContextCollector()],
['project', new ProjectContextCollector()],
['user', new UserContextCollector()],
['session', new SessionContextCollector()]
]);
// Different update frequencies for different context types
this.updateDebouncer = new Map([
['immediate', debounce(this.updateImmediate, 50)], // 50ms - very fast
['document', debounce(this.updateDocument, 200)], // 200ms
['project', debounce(this.updateProject, 5000)], // 5s - slow changing
['user', debounce(this.updateUser, 1000)] // 1s
]);
}
async collectContext(trigger: ContextTrigger): Promise<CopilotContext> {
// Always get fresh immediate context
const immediate = await this.collectors.get('immediate')!.collect(trigger);
// Other contexts can be cached
const [document, project, user, session] = await Promise.all([
this.getCachedOrCollect('document', trigger),
this.getCachedOrCollect('project', trigger),
this.getCachedOrCollect('user', trigger),
this.getCachedOrCollect('session', trigger)
]);
return { immediate, document, project, user, session };
}
private async getCachedOrCollect(
type: ContextType,
trigger: ContextTrigger
): Promise<ContextFragment> {
const cacheKey = this.getCacheKey(type, trigger);
const cached = this.contextCache.get(cacheKey);
if (cached && !this.isStale(cached, type)) {
return cached;
}
const collector = this.collectors.get(type)!;
const context = await collector.collect(trigger);
this.contextCache.set(cacheKey, {
...context,
collectedAt: Date.now()
});
return context;
}
// Subscribe to host application events
subscribeToHostEvents(host: HostApplication): void {
host.on('cursorMove', (position) => {
this.updateDebouncer.get('immediate')!(position);
});
host.on('textChange', (change) => {
this.updateDebouncer.get('immediate')!(change);
this.updateDebouncer.get('document')!(change);
});
host.on('fileOpen', (file) => {
this.updateDebouncer.get('document')!(file);
this.updateDebouncer.get('project')!(file);
});
host.on('selectionChange', (selection) => {
this.updateDebouncer.get('immediate')!(selection);
});
}
}
Code-Specific Context Collection
GIF via GIPHY
class CodeContextCollector implements ContextSourceCollector {
private parser: Parser;
private analyzer: SemanticAnalyzer;
async collect(trigger: ContextTrigger): Promise<CodeContext> {
const { document, position } = trigger;
// Parse current file
const ast = await this.parser.parse(document.content, document.language);
// Find enclosing scope
const scope = this.findEnclosingScope(ast, position);
// Get symbols in scope
const symbols = await this.analyzer.getSymbolsInScope(scope);
// Get type information if available
const typeInfo = await this.getTypeContext(document, position);
// Find related code patterns
const patterns = await this.findSimilarPatterns(document, position);
return {
language: document.language,
currentScope: {
type: scope.type,
name: scope.name,
startLine: scope.startLine,
endLine: scope.endLine
},
enclosingFunction: this.getEnclosingFunction(ast, position),
enclosingClass: this.getEnclosingClass(ast, position),
localVariables: symbols.filter(s => s.kind === 'variable'),
availableFunctions: symbols.filter(s => s.kind === 'function'),
imports: this.extractImports(ast),
typeContext: typeInfo,
recentPatterns: patterns,
syntaxContext: this.getSyntaxContext(ast, position)
};
}
private getSyntaxContext(ast: AST, position: Position): SyntaxContext {
const node = this.findNodeAtPosition(ast, position);
return {
nodeType: node.type,
parentType: node.parent?.type,
isInString: this.isInString(node),
isInComment: this.isInComment(node),
isInFunctionCall: this.isInFunctionCall(node),
isInObjectLiteral: this.isInObjectLiteral(node),
expectedType: this.inferExpectedType(node),
completionKind: this.determineCompletionKind(node)
};
}
private determineCompletionKind(node: ASTNode): CompletionKind {
if (this.isPropertyAccess(node)) return 'property';
if (this.isMethodCall(node)) return 'method';
if (this.isImportStatement(node)) return 'module';
if (this.isTypeAnnotation(node)) return 'type';
if (this.isInFunctionArgs(node)) return 'argument';
return 'general';
}
}
Suggestion Engine Architecture
Multi-Tier Suggestion System
interface SuggestionEngine {
// Instant suggestions - pre-computed, cached
getInstantSuggestion(context: CopilotContext): Suggestion | null;
// Fast suggestions - lightweight model, <100ms
getFastSuggestion(context: CopilotContext): Promise<Suggestion | null>;
// Quality suggestions - full model, <500ms
getQualitySuggestion(context: CopilotContext): Promise<Suggestion>;
// Background pre-computation
precomputeSuggestions(context: CopilotContext): void;
}
class TieredSuggestionEngine implements SuggestionEngine {
private instantCache: SuggestionCache;
private fastModel: FastInferenceClient;
private qualityModel: QualityInferenceClient;
private precomputeQueue: PriorityQueue<PrecomputeTask>;
getInstantSuggestion(context: CopilotContext): Suggestion | null {
// Check exact match cache
const cacheKey = this.computeCacheKey(context);
const cached = this.instantCache.get(cacheKey);
if (cached && this.isValidSuggestion(cached, context)) {
return cached;
}
// Check fuzzy match cache
const fuzzyMatch = this.instantCache.findSimilar(context, {
threshold: 0.9,
maxAge: 30000
});
if (fuzzyMatch) {
return this.adaptSuggestion(fuzzyMatch, context);
}
return null;
}
async getFastSuggestion(context: CopilotContext): Promise<Suggestion | null> {
// Check if instant suggestion is good enough
const instant = this.getInstantSuggestion(context);
if (instant && instant.confidence > 0.8) {
return instant;
}
// Fast model inference
const startTime = performance.now();
try {
const suggestion = await this.fastModel.complete({
prompt: this.buildPrompt(context, 'fast'),
maxTokens: 50,
timeout: 100
});
const latency = performance.now() - startTime;
this.metrics.recordLatency('fast', latency);
if (suggestion && suggestion.confidence > 0.6) {
this.instantCache.set(this.computeCacheKey(context), suggestion);
return suggestion;
}
} catch (error) {
if (error.name !== 'TimeoutError') {
this.metrics.recordError('fast', error);
}
}
return instant; // Fall back to instant if available
}
async getQualitySuggestion(context: CopilotContext): Promise<Suggestion> {
// Start with fast suggestion while quality computes
const fastPromise = this.getFastSuggestion(context);
const qualityPromise = this.qualityModel.complete({
prompt: this.buildPrompt(context, 'quality'),
maxTokens: 200,
temperature: 0.2
});
// Race with timeout - prefer quality if fast enough
const result = await Promise.race([
qualityPromise.then(s => ({ type: 'quality' as const, suggestion: s })),
new Promise<{ type: 'timeout' }>((resolve) =>
setTimeout(() => resolve({ type: 'timeout' }), 300)
)
]);
if (result.type === 'quality') {
return result.suggestion;
}
// Quality timed out, return fast result
const fast = await fastPromise;
if (fast) return fast;
// Wait for quality
return qualityPromise;
}
precomputeSuggestions(context: CopilotContext): void {
// Predict likely next contexts
const predictions = this.predictNextContexts(context);
for (const prediction of predictions) {
this.precomputeQueue.enqueue({
context: prediction.context,
priority: prediction.probability,
deadline: Date.now() + 5000
});
}
// Process queue in background
this.processPrecomputeQueue();
}
private predictNextContexts(
context: CopilotContext
): ContextPrediction[] {
const predictions: ContextPrediction[] = [];
// Predict continuation of current line
const lineCompletion = this.predictLineCompletion(context);
if (lineCompletion) {
predictions.push({
context: { ...context, immediate: lineCompletion },
probability: 0.7
});
}
// Predict next line start
const nextLineContexts = this.predictNextLineContexts(context);
predictions.push(...nextLineContexts);
// Predict based on common patterns
const patternPredictions = this.predictFromPatterns(context);
predictions.push(...patternPredictions);
return predictions.sort((a, b) => b.probability - a.probability).slice(0, 5);
}
}
Suggestion Ranking and Filtering
GIF via GIPHY
class SuggestionRanker {
private weights: RankingWeights;
private userModel: UserPreferenceModel;
async rank(
suggestions: Suggestion[],
context: CopilotContext
): Promise<RankedSuggestion[]> {
const scored = await Promise.all(
suggestions.map(async (suggestion) => ({
suggestion,
score: await this.computeScore(suggestion, context)
}))
);
return scored
.filter(s => s.score.total > this.weights.minimumScore)
.sort((a, b) => b.score.total - a.score.total)
.map((s, index) => ({
...s.suggestion,
rank: index + 1,
scoreBreakdown: s.score
}));
}
private async computeScore(
suggestion: Suggestion,
context: CopilotContext
): Promise<SuggestionScore> {
const scores = {
// Model confidence
modelConfidence: suggestion.confidence * this.weights.modelConfidence,
// Contextual relevance
contextRelevance: this.computeContextRelevance(suggestion, context) *
this.weights.contextRelevance,
// Code quality signals
codeQuality: await this.computeCodeQuality(suggestion, context) *
this.weights.codeQuality,
// User preference alignment
userPreference: this.computeUserPreference(suggestion, context) *
this.weights.userPreference,
// Recency penalty (avoid repeating rejected suggestions)
recencyPenalty: this.computeRecencyPenalty(suggestion, context) *
this.weights.recencyPenalty
};
return {
...scores,
total: Object.values(scores).reduce((a, b) => a + b, 0)
};
}
private computeContextRelevance(
suggestion: Suggestion,
context: CopilotContext
): number {
let relevance = 0;
// Check if suggestion uses available symbols
const usedSymbols = this.extractSymbols(suggestion.content);
const availableSymbols = new Set(
context.document.symbols?.map(s => s.name) || []
);
const symbolMatch = usedSymbols.filter(s => availableSymbols.has(s)).length /
Math.max(usedSymbols.length, 1);
relevance += symbolMatch * 0.3;
// Check type compatibility
if (context.immediate.expectedType && suggestion.inferredType) {
const typeMatch = this.isTypeCompatible(
suggestion.inferredType,
context.immediate.expectedType
);
relevance += typeMatch ? 0.3 : 0;
}
// Check naming convention consistency
const conventionMatch = this.checkNamingConvention(
suggestion.content,
context.project.conventions
);
relevance += conventionMatch * 0.2;
// Check import availability
const importCheck = this.checkImportAvailability(
suggestion.content,
context.document.imports,
context.project.dependencies
);
relevance += importCheck * 0.2;
return relevance;
}
private async computeCodeQuality(
suggestion: Suggestion,
context: CopilotContext
): Promise<number> {
const checks = await Promise.all([
this.checkSyntaxValid(suggestion.content, context.document.language),
this.checkNoObviousErrors(suggestion.content),
this.checkComplexityReasonable(suggestion.content),
this.checkSecurityPatterns(suggestion.content)
]);
return checks.reduce((a, b) => a * b, 1);
}
}
UI Integration Patterns
Ghost Text Rendering
// Ghost text shows suggestions inline without user action
class GhostTextRenderer {
private overlay: HTMLElement;
private currentSuggestion: Suggestion | null = null;
private animationFrame: number | null = null;
constructor(private editor: EditorInterface) {
this.overlay = this.createOverlay();
this.editor.container.appendChild(this.overlay);
}
show(suggestion: Suggestion, position: Position): void {
if (this.animationFrame) {
cancelAnimationFrame(this.animationFrame);
}
this.animationFrame = requestAnimationFrame(() => {
this.currentSuggestion = suggestion;
// Calculate pixel position
const coords = this.editor.getCoordinatesAtPosition(position);
// Render ghost text
this.overlay.innerHTML = '';
this.overlay.style.left = `${coords.x}px`;
this.overlay.style.top = `${coords.y}px`;
const ghostElement = document.createElement('span');
ghostElement.className = 'copilot-ghost-text';
ghostElement.textContent = suggestion.content;
ghostElement.style.opacity = '0';
this.overlay.appendChild(ghostElement);
// Fade in animation
requestAnimationFrame(() => {
ghostElement.style.transition = 'opacity 150ms ease-in';
ghostElement.style.opacity = '0.5';
});
// Set up keyboard handlers
this.setupKeyboardHandlers();
});
}
hide(): void {
const ghostElement = this.overlay.querySelector('.copilot-ghost-text');
if (ghostElement) {
ghostElement.style.opacity = '0';
setTimeout(() => {
this.overlay.innerHTML = '';
this.currentSuggestion = null;
}, 150);
}
}
accept(): void {
if (!this.currentSuggestion) return;
const suggestion = this.currentSuggestion;
// Insert text
this.editor.insertText(suggestion.content);
// Track acceptance
this.trackAcceptance(suggestion);
// Clear
this.hide();
}
private setupKeyboardHandlers(): void {
const handler = (event: KeyboardEvent) => {
if (!this.currentSuggestion) {
this.editor.removeEventListener('keydown', handler);
return;
}
// Tab to accept
if (event.key === 'Tab') {
event.preventDefault();
this.accept();
this.editor.removeEventListener('keydown', handler);
return;
}
// Escape to dismiss
if (event.key === 'Escape') {
this.trackRejection(this.currentSuggestion, 'explicit_dismiss');
this.hide();
this.editor.removeEventListener('keydown', handler);
return;
}
// Any other key dismisses implicitly
if (this.isTypingKey(event)) {
this.trackRejection(this.currentSuggestion, 'continued_typing');
this.hide();
this.editor.removeEventListener('keydown', handler);
}
};
this.editor.addEventListener('keydown', handler);
}
}
// React hook for ghost text
function useGhostText(editorRef: RefObject<EditorInterface>) {
const rendererRef = useRef<GhostTextRenderer | null>(null);
const suggestion = useCopilotStore(state => state.currentSuggestion);
const position = useCopilotStore(state => state.suggestionPosition);
useEffect(() => {
if (editorRef.current && !rendererRef.current) {
rendererRef.current = new GhostTextRenderer(editorRef.current);
}
}, [editorRef.current]);
useEffect(() => {
if (!rendererRef.current) return;
if (suggestion && position) {
rendererRef.current.show(suggestion, position);
} else {
rendererRef.current.hide();
}
}, [suggestion, position]);
return {
accept: () => rendererRef.current?.accept(),
dismiss: () => rendererRef.current?.hide()
};
}
Inline Suggestion Panel
// Panel for multi-line suggestions or choices
function InlineSuggestionPanel({
suggestions,
position,
onAccept,
onDismiss,
onCycle
}: InlineSuggestionPanelProps) {
const [selectedIndex, setSelectedIndex] = useState(0);
const panelRef = useRef<HTMLDivElement>(null);
// Position the panel
const panelStyle = useMemo(() => ({
position: 'absolute' as const,
left: position.x,
top: position.y + position.lineHeight,
maxWidth: '600px',
maxHeight: '300px'
}), [position]);
// Keyboard navigation
useEffect(() => {
const handleKeyDown = (e: KeyboardEvent) => {
switch (e.key) {
case 'ArrowDown':
e.preventDefault();
setSelectedIndex(i => (i + 1) % suggestions.length);
break;
case 'ArrowUp':
e.preventDefault();
setSelectedIndex(i => (i - 1 + suggestions.length) % suggestions.length);
break;
case 'Enter':
case 'Tab':
e.preventDefault();
onAccept(suggestions[selectedIndex]);
break;
case 'Escape':
onDismiss();
break;
}
};
window.addEventListener('keydown', handleKeyDown);
return () => window.removeEventListener('keydown', handleKeyDown);
}, [suggestions, selectedIndex, onAccept, onDismiss]);
return (
<div ref={panelRef} className="suggestion-panel" style={panelStyle}>
<div className="panel-header">
<CopilotIcon />
<span className="suggestion-count">
{selectedIndex + 1} of {suggestions.length}
</span>
<KeyboardShortcut keys={['Tab']} label="Accept" />
<KeyboardShortcut keys={['↑', '↓']} label="Navigate" />
</div>
<div className="suggestions-list">
{suggestions.map((suggestion, index) => (
<SuggestionItem
key={suggestion.id}
suggestion={suggestion}
isSelected={index === selectedIndex}
onClick={() => onAccept(suggestion)}
/>
))}
</div>
<div className="panel-footer">
<ConfidenceIndicator confidence={suggestions[selectedIndex].confidence} />
<button onClick={() => onCycle('next')}>
More suggestions
</button>
</div>
</div>
);
}
function SuggestionItem({
suggestion,
isSelected,
onClick
}: SuggestionItemProps) {
return (
<div
className={`suggestion-item ${isSelected ? 'selected' : ''}`}
onClick={onClick}
>
<SyntaxHighlighter
code={suggestion.content}
language={suggestion.language}
/>
{suggestion.explanation && (
<div className="suggestion-explanation">
{suggestion.explanation}
</div>
)}
</div>
);
}
GIF via GIPHY
Command Palette Integration
// Copilot commands in the command palette
class CopilotCommandProvider implements CommandProvider {
getCommands(context: CommandContext): Command[] {
return [
{
id: 'copilot.explain',
label: 'Copilot: Explain selection',
icon: 'brain',
when: context.hasSelection,
execute: () => this.explainSelection(context.selection)
},
{
id: 'copilot.refactor',
label: 'Copilot: Refactor selection',
icon: 'wand',
when: context.hasSelection,
execute: () => this.refactorSelection(context.selection)
},
{
id: 'copilot.generateTest',
label: 'Copilot: Generate test for function',
icon: 'test',
when: context.isInFunction,
execute: () => this.generateTest(context.enclosingFunction)
},
{
id: 'copilot.fixError',
label: 'Copilot: Fix this error',
icon: 'fix',
when: context.hasError,
execute: () => this.fixError(context.currentError)
},
{
id: 'copilot.complete',
label: 'Copilot: Complete code',
icon: 'sparkle',
keybinding: 'Ctrl+Shift+Space',
execute: () => this.triggerCompletion()
},
{
id: 'copilot.chat',
label: 'Copilot: Open chat',
icon: 'chat',
keybinding: 'Ctrl+Shift+C',
execute: () => this.openCopilotChat()
}
];
}
private async explainSelection(selection: Selection): Promise<void> {
const panel = await this.openResultPanel('explanation');
const stream = this.copilotService.explain({
code: selection.text,
language: selection.document.language,
context: await this.contextCollector.collect({
type: 'selection',
selection
})
});
for await (const chunk of stream) {
panel.appendContent(chunk);
}
}
private async refactorSelection(selection: Selection): Promise<void> {
// Show refactoring options
const option = await this.showRefactorOptions([
{ id: 'simplify', label: 'Simplify' },
{ id: 'extractFunction', label: 'Extract to function' },
{ id: 'optimize', label: 'Optimize for performance' },
{ id: 'modernize', label: 'Modernize syntax' },
{ id: 'custom', label: 'Custom instruction...' }
]);
if (!option) return;
const diff = await this.copilotService.refactor({
code: selection.text,
language: selection.document.language,
instruction: option.instruction,
context: await this.contextCollector.collect({
type: 'selection',
selection
})
});
// Show diff preview
const accepted = await this.showDiffPreview(selection, diff);
if (accepted) {
this.editor.applyEdit(diff.edit);
}
}
}
Trigger System Architecture
Intelligent Trigger Detection
interface TriggerConfig {
type: TriggerType;
conditions: TriggerCondition[];
debounce: number;
priority: number;
}
type TriggerType =
| 'cursor_idle' // User stopped typing
| 'line_end' // Reached end of line
| 'statement_end' // Completed a statement
| 'function_signature' // Typing function parameters
| 'comment_start' // Started a comment
| 'error_hover' // Hovering over error
| 'explicit' // User explicitly requested
| 'new_line'; // Started new line after code
class TriggerEngine {
private triggers: TriggerConfig[];
private activeTimers = new Map<string, NodeJS.Timeout>();
private lastTrigger: number = 0;
private minTriggerInterval = 500;
constructor(triggers: TriggerConfig[]) {
this.triggers = triggers.sort((a, b) => b.priority - a.priority);
}
evaluate(event: EditorEvent): TriggerResult | null {
// Rate limiting
if (Date.now() - this.lastTrigger < this.minTriggerInterval) {
return null;
}
for (const trigger of this.triggers) {
if (this.matchesTrigger(event, trigger)) {
return this.createTriggerResult(event, trigger);
}
}
return null;
}
private matchesTrigger(event: EditorEvent, trigger: TriggerConfig): boolean {
return trigger.conditions.every(condition =>
this.evaluateCondition(condition, event)
);
}
private evaluateCondition(
condition: TriggerCondition,
event: EditorEvent
): boolean {
switch (condition.type) {
case 'idle_time':
return event.timeSinceLastKeypress >= condition.value;
case 'cursor_position':
return this.matchCursorPosition(event.cursor, condition);
case 'syntax_context':
return this.matchSyntaxContext(event.syntaxContext, condition);
case 'line_content':
return condition.pattern.test(event.currentLine);
case 'document_state':
return this.matchDocumentState(event.document, condition);
case 'user_preference':
return this.checkUserPreference(condition);
default:
return false;
}
}
private matchCursorPosition(
cursor: CursorPosition,
condition: TriggerCondition
): boolean {
switch (condition.position) {
case 'end_of_line':
return cursor.column === cursor.lineLength;
case 'end_of_statement':
return this.isEndOfStatement(cursor);
case 'inside_function_call':
return cursor.syntaxContext.isInFunctionCall;
case 'after_operator':
return this.isAfterOperator(cursor);
default:
return false;
}
}
}
// Default trigger configurations
const defaultTriggers: TriggerConfig[] = [
{
type: 'cursor_idle',
conditions: [
{ type: 'idle_time', value: 750 },
{ type: 'syntax_context', notIn: ['string', 'comment'] }
],
debounce: 100,
priority: 10
},
{
type: 'line_end',
conditions: [
{ type: 'cursor_position', position: 'end_of_line' },
{ type: 'line_content', pattern: /^.+[^,{(\[]$/ }
],
debounce: 300,
priority: 20
},
{
type: 'function_signature',
conditions: [
{ type: 'syntax_context', in: ['function_parameter'] },
{ type: 'idle_time', value: 500 }
],
debounce: 200,
priority: 30
},
{
type: 'comment_start',
conditions: [
{ type: 'line_content', pattern: /^\s*\/\/\s*$/ }
],
debounce: 100,
priority: 25
},
{
type: 'new_line',
conditions: [
{ type: 'cursor_position', position: 'start_of_line' },
{ type: 'document_state', previousLineHasCode: true }
],
debounce: 400,
priority: 15
}
];
Adaptive Trigger Tuning
GIF via GIPHY
class AdaptiveTriggerManager {
private userBehavior: UserBehaviorModel;
private triggerEffectiveness = new Map<TriggerType, TriggerStats>();
constructor(private baseEngine: TriggerEngine) {
this.userBehavior = new UserBehaviorModel();
}
evaluate(event: EditorEvent): TriggerResult | null {
// Adjust triggers based on user behavior
const adjustedTriggers = this.adjustTriggers(event);
const result = this.baseEngine.evaluate(event);
if (result) {
this.trackTrigger(result.type);
}
return result;
}
recordSuggestionOutcome(
trigger: TriggerType,
outcome: 'accepted' | 'rejected' | 'ignored'
): void {
const stats = this.triggerEffectiveness.get(trigger) || {
accepted: 0,
rejected: 0,
ignored: 0
};
stats[outcome]++;
this.triggerEffectiveness.set(trigger, stats);
// Update user behavior model
this.userBehavior.recordOutcome(trigger, outcome);
}
private adjustTriggers(event: EditorEvent): TriggerConfig[] {
const userPrefs = this.userBehavior.getPreferences();
return this.baseEngine.triggers.map(trigger => {
const effectiveness = this.getEffectiveness(trigger.type);
const userPref = userPrefs[trigger.type];
// Adjust debounce based on typing speed
const adjustedDebounce = trigger.debounce *
(userPrefs.typingSpeed / 100);
// Adjust priority based on effectiveness
const adjustedPriority = trigger.priority *
(effectiveness > 0.3 ? 1 : 0.5);
// Disable if user consistently rejects
if (userPref?.disabled || effectiveness < 0.1) {
return { ...trigger, conditions: [{ type: 'never' }] };
}
return {
...trigger,
debounce: adjustedDebounce,
priority: adjustedPriority
};
});
}
private getEffectiveness(type: TriggerType): number {
const stats = this.triggerEffectiveness.get(type);
if (!stats) return 0.5;
const total = stats.accepted + stats.rejected + stats.ignored;
if (total < 10) return 0.5; // Not enough data
return stats.accepted / total;
}
}
Caching and Performance
Multi-Level Suggestion Cache
interface SuggestionCache {
// L1: In-memory, exact match
getExact(key: string): CachedSuggestion | null;
// L2: In-memory, fuzzy match
getSimilar(context: CopilotContext, threshold: number): CachedSuggestion | null;
// L3: Persistent, semantic match
getSemantic(embedding: number[]): Promise<CachedSuggestion | null>;
// Pre-computation storage
storePrecomputed(context: CopilotContext, suggestion: Suggestion): void;
}
class TieredSuggestionCache implements SuggestionCache {
private l1Cache: LRUCache<string, CachedSuggestion>;
private l2Cache: FuzzyMatchCache;
private l3Store: VectorStore;
constructor(config: CacheConfig) {
this.l1Cache = new LRUCache({
max: config.l1MaxEntries,
ttl: config.l1TtlMs
});
this.l2Cache = new FuzzyMatchCache({
maxEntries: config.l2MaxEntries,
similarityThreshold: config.l2SimilarityThreshold
});
this.l3Store = new VectorStore(config.l3Config);
}
getExact(key: string): CachedSuggestion | null {
return this.l1Cache.get(key) || null;
}
getSimilar(
context: CopilotContext,
threshold: number
): CachedSuggestion | null {
// Build context signature for fuzzy matching
const signature = this.buildContextSignature(context);
const match = this.l2Cache.findMatch(signature, threshold);
if (match) {
// Promote to L1 on hit
this.l1Cache.set(this.buildExactKey(context), match);
}
return match;
}
async getSemantic(embedding: number[]): Promise<CachedSuggestion | null> {
const results = await this.l3Store.search(embedding, {
limit: 1,
minScore: 0.9
});
if (results.length > 0) {
const cached = results[0].payload as CachedSuggestion;
// Promote to L2
this.l2Cache.set(cached.contextSignature, cached);
return cached;
}
return null;
}
storePrecomputed(context: CopilotContext, suggestion: Suggestion): void {
const exactKey = this.buildExactKey(context);
const signature = this.buildContextSignature(context);
const cached: CachedSuggestion = {
suggestion,
contextSignature: signature,
createdAt: Date.now(),
accessCount: 0
};
// Store in L1
this.l1Cache.set(exactKey, cached);
// Store in L2
this.l2Cache.set(signature, cached);
// Async store in L3 for long-term
this.storeInL3(context, cached).catch(console.error);
}
private buildContextSignature(context: CopilotContext): string {
return [
context.immediate.currentLine,
context.immediate.syntaxContext?.nodeType,
context.document.type,
context.document.language
].join('|');
}
private async storeInL3(
context: CopilotContext,
cached: CachedSuggestion
): Promise<void> {
const embedding = await this.embedder.embed(
this.buildContextSignature(context)
);
await this.l3Store.insert({
id: generateId(),
embedding,
payload: cached
});
}
}
Speculative Pre-fetching
GIF via GIPHY
class SpeculativePrefetcher {
private prefetchQueue: PriorityQueue<PrefetchTask>;
private inFlightRequests = new Map<string, Promise<Suggestion>>();
private maxConcurrent = 3;
constructor(
private suggestionEngine: SuggestionEngine,
private cache: SuggestionCache
) {}
onContextChange(context: CopilotContext): void {
// Predict likely next contexts
const predictions = this.predictNextContexts(context);
// Queue prefetch tasks
for (const prediction of predictions) {
if (prediction.probability > 0.3) {
this.enqueuePrefetch(prediction);
}
}
// Process queue
this.processQueue();
}
private predictNextContexts(
current: CopilotContext
): ContextPrediction[] {
const predictions: ContextPrediction[] = [];
// Predict typing continuation
if (current.immediate.currentLine.length > 0) {
const continuations = this.predictContinuations(
current.immediate.currentLine
);
for (const cont of continuations) {
predictions.push({
context: {
...current,
immediate: {
...current.immediate,
currentLine: current.immediate.currentLine + cont.text
}
},
probability: cont.probability
});
}
}
// Predict new line scenarios
const newLineContext = this.predictNewLineContext(current);
if (newLineContext) {
predictions.push(newLineContext);
}
// Predict based on document patterns
const patternPredictions = this.predictFromPatterns(current);
predictions.push(...patternPredictions);
return predictions;
}
private async processQueue(): Promise<void> {
while (
this.prefetchQueue.size() > 0 &&
this.inFlightRequests.size < this.maxConcurrent
) {
const task = this.prefetchQueue.dequeue();
if (!task) break;
// Skip if already cached
const cacheKey = this.buildCacheKey(task.context);
if (this.cache.getExact(cacheKey)) {
continue;
}
// Skip if already in flight
if (this.inFlightRequests.has(cacheKey)) {
continue;
}
// Start prefetch
const promise = this.prefetch(task);
this.inFlightRequests.set(cacheKey, promise);
promise.finally(() => {
this.inFlightRequests.delete(cacheKey);
this.processQueue();
});
}
}
private async prefetch(task: PrefetchTask): Promise<Suggestion | null> {
try {
const suggestion = await this.suggestionEngine.getFastSuggestion(
task.context
);
if (suggestion) {
this.cache.storePrecomputed(task.context, suggestion);
}
return suggestion;
} catch (error) {
// Prefetch failures are non-critical
console.debug('Prefetch failed:', error);
return null;
}
}
}
Learning and Personalization
User Preference Learning
interface UserPreferenceModel {
// Code style preferences
codeStyle: {
preferredNamingConvention: NamingConvention;
preferredQuoteStyle: 'single' | 'double';
preferredIndentation: number;
bracketStyle: 'same-line' | 'new-line';
};
// Suggestion preferences
suggestionPreferences: {
preferredLength: 'short' | 'medium' | 'long';
includeComments: boolean;
includeTypes: boolean;
verbosity: number;
};
// Trigger preferences
triggerPreferences: Map<TriggerType, TriggerPreference>;
// Learning signals
acceptancePatterns: AcceptancePattern[];
rejectionPatterns: RejectionPattern[];
}
class UserLearningEngine {
private model: UserPreferenceModel;
private eventBuffer: UserEvent[];
private updateDebounce = debounce(this.updateModel.bind(this), 5000);
recordAcceptance(
suggestion: Suggestion,
context: CopilotContext,
modifications: string | null
): void {
const event: AcceptanceEvent = {
type: 'acceptance',
timestamp: Date.now(),
suggestion,
context,
wasModified: !!modifications,
modifications
};
this.eventBuffer.push(event);
this.updateDebounce();
// Extract immediate learnings
if (!modifications) {
// Full acceptance - strong signal
this.reinforcePattern(suggestion, context, 1.0);
} else {
// Modified acceptance - learn from modifications
this.learnFromModification(suggestion, modifications, context);
}
}
recordRejection(
suggestion: Suggestion,
context: CopilotContext,
reason: RejectionReason
): void {
const event: RejectionEvent = {
type: 'rejection',
timestamp: Date.now(),
suggestion,
context,
reason
};
this.eventBuffer.push(event);
this.updateDebounce();
// Learn from rejection
if (reason === 'wrong_suggestion') {
this.penalizePattern(suggestion, context, 0.8);
} else if (reason === 'timing') {
this.adjustTriggerTiming(context);
}
}
private learnFromModification(
original: Suggestion,
modified: string,
context: CopilotContext
): void {
// Analyze the diff
const diff = this.computeDiff(original.content, modified);
// Learn naming preferences
const namingChanges = this.extractNamingChanges(diff);
for (const change of namingChanges) {
this.updateNamingPreference(change.original, change.modified);
}
// Learn style preferences
const styleChanges = this.extractStyleChanges(diff);
for (const change of styleChanges) {
this.updateStylePreference(change);
}
// Learn verbosity preferences
const lengthChange = modified.length / original.content.length;
this.updateVerbosityPreference(lengthChange);
}
getSuggestionAdjustments(
suggestion: Suggestion,
context: CopilotContext
): SuggestionAdjustments {
return {
// Apply learned naming conventions
renameIdentifiers: this.getRenamings(suggestion, context),
// Apply style preferences
reformatting: this.getReformatting(suggestion),
// Adjust verbosity
verbosityAdjustment: this.getVerbosityAdjustment(suggestion),
// Apply type annotation preferences
typeAnnotations: this.getTypeAnnotationAdjustment(suggestion)
};
}
private async updateModel(): Promise<void> {
if (this.eventBuffer.length === 0) return;
const events = [...this.eventBuffer];
this.eventBuffer = [];
// Batch process events
const patterns = this.extractPatterns(events);
// Update model
for (const pattern of patterns) {
this.model.acceptancePatterns.push(pattern);
}
// Prune old patterns
this.pruneOldPatterns();
// Persist model
await this.persistModel();
}
}
GIF via GIPHY
Observability and Metrics
Copilot-Specific Metrics
interface CopilotMetrics {
// Suggestion metrics
suggestionsShown: Counter;
suggestionLatency: Histogram;
suggestionAcceptanceRate: Gauge;
// Cache metrics
cacheHitRate: Gauge;
prefetchHitRate: Gauge;
// Quality metrics
suggestionQualityScore: Histogram;
userModificationRate: Gauge;
// Trigger metrics
triggersByType: Counter;
triggerEffectiveness: Gauge;
// Performance metrics
contextCollectionTime: Histogram;
renderTime: Histogram;
}
class CopilotTelemetry {
private metrics: CopilotMetrics;
trackSuggestionCycle(cycle: SuggestionCycle): void {
// Latency tracking
this.metrics.suggestionLatency.record(cycle.totalLatency, {
tier: cycle.suggestionTier,
cacheHit: cycle.cacheHit
});
// Show tracking
this.metrics.suggestionsShown.add(1, {
triggerType: cycle.triggerType,
suggestionType: cycle.suggestionType
});
// Outcome tracking (when available)
if (cycle.outcome) {
this.trackOutcome(cycle);
}
}
private trackOutcome(cycle: SuggestionCycle): void {
const labels = {
triggerType: cycle.triggerType,
suggestionType: cycle.suggestionType
};
switch (cycle.outcome) {
case 'accepted':
this.metrics.suggestionAcceptanceRate.set(
this.calculateAcceptanceRate(),
labels
);
break;
case 'modified':
this.metrics.userModificationRate.set(
this.calculateModificationRate(),
labels
);
break;
}
}
generateDashboard(): DashboardConfig {
return {
panels: [
{
title: 'Suggestion Performance',
metrics: [
'copilot_suggestion_latency_p50',
'copilot_suggestion_latency_p99',
'copilot_suggestions_shown_rate',
'copilot_suggestion_acceptance_rate'
]
},
{
title: 'Cache Effectiveness',
metrics: [
'copilot_cache_hit_rate',
'copilot_prefetch_hit_rate',
'copilot_l1_cache_size',
'copilot_l2_cache_size'
]
},
{
title: 'Trigger Analysis',
metrics: [
'copilot_triggers_by_type',
'copilot_trigger_effectiveness',
'copilot_trigger_false_positive_rate'
]
},
{
title: 'User Experience',
metrics: [
'copilot_time_saved_estimate',
'copilot_user_modification_rate',
'copilot_rejection_rate_by_reason'
]
}
],
alerts: [
{
name: 'High Suggestion Latency',
condition: 'copilot_suggestion_latency_p99 > 500',
severity: 'warning'
},
{
name: 'Low Acceptance Rate',
condition: 'copilot_suggestion_acceptance_rate < 0.15',
severity: 'warning'
},
{
name: 'Cache Degradation',
condition: 'copilot_cache_hit_rate < 0.3',
severity: 'critical'
}
]
};
}
}
GIF via GIPHY
Production Incidents & Lessons
Incident 1: Suggestion Storm
Symptoms: Users reported UI freezing with rapid suggestion flashing. CPU usage spiked to 100%.
Root Cause: Trigger debouncing was too aggressive for fast typists. Each keystroke triggered context collection, which triggered prefetching, which triggered rendering.
// Before: Simple debounce
const triggerSuggestion = debounce(async () => {
const context = await collectContext(); // Expensive
const suggestion = await getSuggestion(context);
renderSuggestion(suggestion);
}, 100);
// After: Layered debouncing with cancellation
class SuggestionController {
private contextCollectionId = 0;
private suggestionRequestId = 0;
onInput = debounce(async () => {
// Cancel any in-flight work
const currentContextId = ++this.contextCollectionId;
// Lightweight trigger check first
if (!this.shouldTrigger()) return;
// Collect context with cancellation check
const context = await this.collectContext();
if (currentContextId !== this.contextCollectionId) return;
const currentSuggestionId = ++this.suggestionRequestId;
// Get suggestion with cancellation check
const suggestion = await this.getSuggestion(context);
if (currentSuggestionId !== this.suggestionRequestId) return;
// Only render if still relevant
if (this.isStillRelevant(context)) {
this.renderSuggestion(suggestion);
}
}, 150);
private shouldTrigger(): boolean {
// Quick checks before expensive operations
if (this.isUserActivelyTyping()) return false;
if (this.recentlyDismissed()) return false;
return true;
}
}
Incident 2: Context Explosion
Symptoms: Memory usage grew unbounded over long sessions. Eventually crashed browser tabs.
Root Cause: Context collector was storing full document snapshots on every change without pruning.
GIF via GIPHY
// Before: Unbounded history
class ContextCollector {
private documentHistory: DocumentSnapshot[] = [];
onDocumentChange(doc: Document) {
this.documentHistory.push({
content: doc.content,
timestamp: Date.now()
}); // Grows forever
}
}
// After: Bounded sliding window with compression
class ContextCollector {
private documentHistory: CompressedSnapshot[] = [];
private maxSnapshots = 50;
private maxMemoryMB = 10;
onDocumentChange(doc: Document) {
// Compute delta, not full snapshot
const delta = this.computeDelta(doc);
// Add with bounds checking
this.documentHistory.push({
delta,
timestamp: Date.now(),
size: delta.length
});
// Prune if needed
this.pruneIfNeeded();
}
private pruneIfNeeded() {
// Prune by count
while (this.documentHistory.length > this.maxSnapshots) {
this.documentHistory.shift();
}
// Prune by memory
while (this.getTotalSize() > this.maxMemoryMB * 1024 * 1024) {
this.documentHistory.shift();
}
}
}
Incident 3: Suggestion Hallucination
Symptoms: Copilot suggested code using non-existent APIs and undefined variables.
Root Cause: Context window was too small, cutting off import statements and variable declarations.
// Before: Fixed context window
function buildPrompt(context: CopilotContext): string {
return `
${context.immediate.surroundingLines.slice(-10).join('\n')}
// cursor here
`;
}
// After: Semantic context selection
function buildPrompt(context: CopilotContext): string {
const sections: string[] = [];
// Always include imports
sections.push(context.document.imports.join('\n'));
// Include relevant type definitions
const relevantTypes = selectRelevantTypes(context);
sections.push(relevantTypes.join('\n'));
// Include enclosing scope
sections.push(context.immediate.enclosingScope);
// Include immediate context
sections.push(context.immediate.surroundingLines.join('\n'));
// Ensure we're under token limit
return truncateToTokenLimit(sections.join('\n\n'), 2000);
}
function selectRelevantTypes(context: CopilotContext): string[] {
// Use AST analysis to find types used in current scope
const usedTypes = extractUsedTypes(context.immediate.enclosingScope);
// Get their definitions
return usedTypes
.map(t => context.document.typeDefinitions.get(t))
.filter(Boolean);
}
Tradeoffs & Engineering Decisions
Decision: Proactive vs On-Demand Suggestions
| Factor | Proactive | On-Demand |
|---|---|---|
| User discovery | High | Low |
| Annoyance potential | High | None |
| Latency requirements | Strict (<200ms) | Relaxed |
| Resource usage | Higher | Lower |
| Acceptance rate | Lower | Higher |
Decision: Hybrid approach. Proactive suggestions for high-confidence, low-intrusiveness scenarios (ghost text at line end). On-demand for more complex suggestions (command palette, explicit triggers).
Decision: Client-side vs Server-side Inference
| Factor | Client-side | Server-side |
|---|---|---|
| Latency | Lower | Higher |
| Model quality | Limited | Best |
| Privacy | Better | Requires trust |
| Offline support | Yes | No |
| Resource constraints | Device-limited | Scalable |
Decision: Tiered approach. Small, fast model client-side for instant suggestions. Quality model server-side for complex suggestions. Fall back gracefully when offline.
Decision: Full Context vs Minimal Context
GIF via GIPHY
Full Context:
- Pros: Better suggestion quality, understands more
- Cons: Higher latency, token costs, privacy concerns
Minimal Context:
- Pros: Faster, cheaper, more private
- Cons: Lower quality, more hallucinations
Decision: Adaptive context sizing. Start with minimal context for fast suggestions. Expand context for quality suggestions. Always include semantic essentials (imports, types, scope).
Conclusion
Building AI copilots requires a different architectural mindset than standalone AI applications. The copilot must be deeply integrated, contextually aware, predictively intelligent, and above all—unobtrusive.
Key architectural principles:
GIF via GIPHY
- Speculation over reaction: Pre-compute and cache aggressively. Users won't wait.
- Context is king: Invest heavily in multi-source, semantic context collection.
- Tiered suggestions: Match suggestion quality to latency requirements.
- Learn continuously: Personalization dramatically improves acceptance rates.
- Measure obsessively: Acceptance rate is the north star metric.
The copilot pattern represents the future of human-AI collaboration in creative tools. As models improve and latencies decrease, the line between user intent and AI assistance will continue to blur—making thoughtful architecture even more critical.
What did you think?