Designing Scalable Agentic UI: Architecture, Safety, and Control
Introduction
Traditional UI systems are reactive—they wait for user input and respond with deterministic outcomes. Agentic UI systems invert this relationship. The AI doesn't just respond to commands; it plans, executes multi-step workflows, interacts with external systems, makes decisions, and drives the interface toward goals.
Building agentic UI presents unique architectural challenges that traditional frontend patterns don't address. When your interface hosts an autonomous agent that can execute code, call APIs, modify state, and chain multiple operations—the lines between "user action" and "system action" blur entirely. The UI must simultaneously give users visibility into agent reasoning, control over execution, ability to intervene mid-workflow, and confidence that the agent won't cause unintended consequences.
GIF via GIPHY
Production agentic systems at companies like Replit, Cursor, and Devin demonstrate that this architecture is viable at scale—but the implementation details reveal hard-won lessons about safety boundaries, execution sandboxing, state management, and the fundamental tension between agent autonomy and user control.
Scale Context
Modern agentic UI systems operate under demanding constraints:
| Metric | Production Scale |
|---|---|
| Concurrent agent sessions | 10K-100K |
| Tool calls per session | 5-500 |
| Average workflow steps | 3-50 |
| Tool execution P99 latency | 100ms-30s |
| Context window per agent | 16K-200K tokens |
| State checkpoints per session | 10-100 |
| Rollback frequency | 5-15% of sessions |
| Human intervention rate | 10-30% of workflows |
| Concurrent tool executions | 1-10 per agent |
| Memory footprint per session | 10-100 MB |
GIF via GIPHY
The challenge: orchestrating autonomous multi-step execution while maintaining user trust, system safety, and responsive UI feedback throughout workflows that may take minutes to complete.
High-Level Architecture
┌─────────────────────────────────────────────────────────────────────────────┐
│ Agentic UI Frontend │
├─────────────────────────────────────────────────────────────────────────────┤
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Workflow │ │ Tool │ │ Agent │ │ Control │ │
│ │ Viewer │ │ Visualizer │ │ Thoughts │ │ Panel │ │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │ │
│ ┌──────┴─────────────────┴─────────────────┴─────────────────┴──────┐ │
│ │ Agent Orchestration Layer │ │
│ │ ┌───────────┐ ┌───────────┐ ┌───────────┐ ┌───────────┐ │ │
│ │ │ Plan │ │ Tool │ │ State │ │ Rollback │ │ │
│ │ │ Manager │ │ Executor │ │ Manager │ │ Manager │ │ │
│ │ └───────────┘ └───────────┘ └───────────┘ └───────────┘ │ │
│ └────────────────────────────┬──────────────────────────────────────┘ │
│ │ │
│ ┌────────────────────────────┴──────────────────────────────────────┐ │
│ │ Safety & Control Layer │ │
│ │ ┌───────────┐ ┌───────────┐ ┌───────────┐ ┌───────────┐ │ │
│ │ │ Permission│ │ Sandbox │ │ Rate │ │ Approval │ │ │
│ │ │ Manager │ │ Manager │ │ Limiter │ │ Queue │ │ │
│ │ └───────────┘ └───────────┘ └───────────┘ └───────────┘ │ │
│ └────────────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────────────┘
│
┌───────────────┼───────────────┐
▼ ▼ ▼
┌───────────┐ ┌───────────┐ ┌───────────┐
│ Agent │ │ Tool │ │ State │
│ BFF │ │ Gateway │ │ Service │
└─────┬─────┘ └─────┬─────┘ └───────────┘
│ │
▼ ▼
┌─────────────────────────────────────────┐
│ Agent Execution Layer │
│ ┌─────────┐ ┌─────────┐ ┌─────────┐ │
│ │ LLM │ │ Tool │ │ Sandbox │ │
│ │ Planner │ │ Runtime │ │ Envs │ │
│ └─────────┘ └─────────┘ └─────────┘ │
└─────────────────────────────────────────┘
Agent Execution Loop
User Goal → Plan Generation → Step Selection →
Tool Execution → Result Analysis → State Update →
Continue/Complete Decision → Next Step or Finish
GIF via GIPHY
The fundamental insight: agentic systems implement a perceive-plan-act loop that runs continuously until goal completion or human intervention. The frontend must visualize, control, and potentially interrupt this loop at any point.
Agent Loop Architecture
Core Agent Loop Implementation
interface AgentLoop {
sessionId: string;
goal: string;
context: AgentContext;
status: 'planning' | 'executing' | 'waiting' | 'complete' | 'failed' | 'paused';
plan: AgentPlan | null;
currentStep: number;
executionHistory: ExecutionStep[];
// Control
pause(): Promise<void>;
resume(): Promise<void>;
cancel(): Promise<void>;
rollback(toStep: number): Promise<void>;
}
interface AgentPlan {
id: string;
goal: string;
steps: PlanStep[];
confidence: number;
estimatedDuration: number;
requiredPermissions: Permission[];
}
interface PlanStep {
id: string;
description: string;
tool: string;
parameters: Record<string, unknown>;
dependencies: string[];
estimatedDuration: number;
riskLevel: 'low' | 'medium' | 'high';
requiresApproval: boolean;
}
class AgentOrchestrator {
private activeAgents = new Map<string, AgentLoop>();
private eventEmitter = new EventEmitter();
async startAgent(goal: string, context: AgentContext): Promise<AgentLoop> {
const sessionId = generateId();
const agent: AgentLoop = {
sessionId,
goal,
context,
status: 'planning',
plan: null,
currentStep: 0,
executionHistory: [],
pause: () => this.pauseAgent(sessionId),
resume: () => this.resumeAgent(sessionId),
cancel: () => this.cancelAgent(sessionId),
rollback: (step) => this.rollbackAgent(sessionId, step)
};
this.activeAgents.set(sessionId, agent);
this.emit('agent:started', agent);
// Start the agent loop
this.runAgentLoop(agent).catch(error => {
agent.status = 'failed';
this.emit('agent:failed', { agent, error });
});
return agent;
}
private async runAgentLoop(agent: AgentLoop): Promise<void> {
// Phase 1: Planning
agent.status = 'planning';
this.emit('agent:planning', agent);
agent.plan = await this.generatePlan(agent);
this.emit('agent:planned', { agent, plan: agent.plan });
// Check if plan requires upfront approval
if (this.requiresApproval(agent.plan)) {
agent.status = 'waiting';
this.emit('agent:awaiting_approval', { agent, plan: agent.plan });
await this.waitForApproval(agent.sessionId);
}
// Phase 2: Execution
agent.status = 'executing';
while (agent.currentStep < agent.plan.steps.length) {
// Check for pause/cancel
if (agent.status === 'paused') {
await this.waitForResume(agent.sessionId);
}
if (agent.status === 'failed') {
break;
}
const step = agent.plan.steps[agent.currentStep];
this.emit('agent:step_started', { agent, step });
// Step-level approval for high-risk actions
if (step.requiresApproval) {
agent.status = 'waiting';
this.emit('agent:step_awaiting_approval', { agent, step });
await this.waitForStepApproval(agent.sessionId, step.id);
agent.status = 'executing';
}
// Execute the step
const result = await this.executeStep(agent, step);
agent.executionHistory.push(result);
if (result.status === 'failed') {
// Attempt recovery or fail
const recovered = await this.attemptRecovery(agent, step, result);
if (!recovered) {
agent.status = 'failed';
this.emit('agent:step_failed', { agent, step, result });
break;
}
}
this.emit('agent:step_completed', { agent, step, result });
agent.currentStep++;
// Re-plan if needed based on execution results
if (this.shouldReplan(agent, result)) {
agent.status = 'planning';
agent.plan = await this.generatePlan(agent, result);
this.emit('agent:replanned', { agent, plan: agent.plan });
agent.currentStep = 0;
agent.status = 'executing';
}
}
if (agent.status !== 'failed') {
agent.status = 'complete';
this.emit('agent:completed', agent);
}
}
private async generatePlan(
agent: AgentLoop,
previousResult?: ExecutionResult
): Promise<AgentPlan> {
const planningPrompt = this.buildPlanningPrompt(agent, previousResult);
const response = await this.llmClient.complete({
messages: [
{ role: 'system', content: this.planningSystemPrompt },
{ role: 'user', content: planningPrompt }
],
tools: [this.planningTool],
toolChoice: { type: 'tool', name: 'create_plan' }
});
return this.parsePlan(response.toolCalls[0].arguments);
}
private async executeStep(
agent: AgentLoop,
step: PlanStep
): Promise<ExecutionResult> {
const startTime = Date.now();
try {
// Validate permissions
const hasPermission = await this.permissionManager.check(
agent.context.userId,
step.tool,
step.parameters
);
if (!hasPermission) {
return {
stepId: step.id,
status: 'failed',
error: 'Permission denied',
duration: Date.now() - startTime
};
}
// Execute in sandbox
const result = await this.toolExecutor.execute(
step.tool,
step.parameters,
{
timeout: step.estimatedDuration * 2,
sandbox: agent.context.sandboxConfig,
checkpoints: true
}
);
return {
stepId: step.id,
status: 'success',
result,
duration: Date.now() - startTime,
checkpoint: await this.createCheckpoint(agent)
};
} catch (error) {
return {
stepId: step.id,
status: 'failed',
error: error.message,
duration: Date.now() - startTime
};
}
}
}
Tool Registration and Execution
GIF via GIPHY
interface Tool {
name: string;
description: string;
parameters: JSONSchema;
permissions: Permission[];
riskLevel: 'low' | 'medium' | 'high';
timeout: number;
execute(params: unknown, context: ExecutionContext): Promise<ToolResult>;
validate(params: unknown): ValidationResult;
rollback?(params: unknown, result: ToolResult): Promise<void>;
}
class ToolRegistry {
private tools = new Map<string, Tool>();
private executionLimits = new Map<string, RateLimit>();
register(tool: Tool): void {
this.tools.set(tool.name, tool);
// Set default rate limits based on risk
this.executionLimits.set(tool.name, {
low: { requests: 100, window: 60000 },
medium: { requests: 20, window: 60000 },
high: { requests: 5, window: 60000 }
}[tool.riskLevel]);
}
getToolsForLLM(): ToolDefinition[] {
return Array.from(this.tools.values()).map(tool => ({
type: 'function',
function: {
name: tool.name,
description: tool.description,
parameters: tool.parameters
}
}));
}
}
// Example tools for a coding agent
const fileSystemTools: Tool[] = [
{
name: 'read_file',
description: 'Read the contents of a file',
parameters: {
type: 'object',
properties: {
path: { type: 'string', description: 'File path to read' }
},
required: ['path']
},
permissions: ['fs:read'],
riskLevel: 'low',
timeout: 5000,
async execute(params, context) {
const { path } = params as { path: string };
// Validate path is within allowed directories
if (!context.sandbox.isPathAllowed(path)) {
throw new Error(`Access denied: ${path}`);
}
const content = await context.sandbox.readFile(path);
return { type: 'text', content };
},
validate(params) {
const { path } = params as { path: string };
if (!path || typeof path !== 'string') {
return { valid: false, errors: ['Invalid path'] };
}
return { valid: true };
}
},
{
name: 'write_file',
description: 'Write content to a file',
parameters: {
type: 'object',
properties: {
path: { type: 'string' },
content: { type: 'string' }
},
required: ['path', 'content']
},
permissions: ['fs:write'],
riskLevel: 'medium',
timeout: 10000,
async execute(params, context) {
const { path, content } = params as { path: string; content: string };
// Create backup for rollback
const backup = await context.sandbox.readFile(path).catch(() => null);
await context.sandbox.writeFile(path, content);
return {
type: 'success',
backup,
bytesWritten: content.length
};
},
async rollback(params, result) {
const { path } = params as { path: string };
if (result.backup) {
await this.execute({ path, content: result.backup }, context);
} else {
await context.sandbox.deleteFile(path);
}
}
},
{
name: 'execute_command',
description: 'Execute a shell command',
parameters: {
type: 'object',
properties: {
command: { type: 'string' },
workingDirectory: { type: 'string' }
},
required: ['command']
},
permissions: ['shell:execute'],
riskLevel: 'high',
timeout: 60000,
async execute(params, context) {
const { command, workingDirectory } = params as {
command: string;
workingDirectory?: string
};
// Command validation
const sanitized = this.sanitizeCommand(command);
const result = await context.sandbox.exec(sanitized, {
cwd: workingDirectory,
timeout: this.timeout,
maxBuffer: 1024 * 1024 // 1MB
});
return {
type: 'command_result',
stdout: result.stdout,
stderr: result.stderr,
exitCode: result.exitCode
};
},
validate(params) {
const { command } = params as { command: string };
// Block dangerous patterns
const blockedPatterns = [
/rm\s+-rf\s+[\/~]/,
/>\s*\/dev\/sd/,
/mkfs/,
/dd\s+if=/,
/:(){ :|:& };:/ // Fork bomb
];
for (const pattern of blockedPatterns) {
if (pattern.test(command)) {
return {
valid: false,
errors: ['Command contains blocked pattern']
};
}
}
return { valid: true };
}
}
];
State Management for Agentic Systems
Checkpoint and Rollback Architecture
interface AgentState {
sessionId: string;
checkpoints: Checkpoint[];
currentCheckpointId: string | null;
// Derived state
fileSystem: FileSystemState;
environment: EnvironmentState;
memory: AgentMemory;
}
interface Checkpoint {
id: string;
timestamp: number;
stepId: string;
stepDescription: string;
// State snapshot
fileSystemDelta: FileSystemDelta;
environmentDelta: EnvironmentDelta;
memorySnapshot: AgentMemory;
// Rollback capability
canRollback: boolean;
rollbackActions: RollbackAction[];
}
class CheckpointManager {
private checkpoints = new Map<string, Checkpoint[]>();
private maxCheckpointsPerSession = 50;
async createCheckpoint(
sessionId: string,
step: ExecutionResult
): Promise<Checkpoint> {
const previousCheckpoint = this.getLatestCheckpoint(sessionId);
// Capture state deltas
const fileSystemDelta = await this.captureFileSystemDelta(
sessionId,
previousCheckpoint
);
const environmentDelta = await this.captureEnvironmentDelta(
sessionId,
previousCheckpoint
);
const checkpoint: Checkpoint = {
id: generateId(),
timestamp: Date.now(),
stepId: step.stepId,
stepDescription: step.description,
fileSystemDelta,
environmentDelta,
memorySnapshot: await this.captureMemory(sessionId),
canRollback: this.canRollback(step),
rollbackActions: this.buildRollbackActions(step)
};
// Store checkpoint
const sessionCheckpoints = this.checkpoints.get(sessionId) || [];
sessionCheckpoints.push(checkpoint);
// Prune old checkpoints if needed
if (sessionCheckpoints.length > this.maxCheckpointsPerSession) {
this.pruneCheckpoints(sessionId, sessionCheckpoints);
}
this.checkpoints.set(sessionId, sessionCheckpoints);
return checkpoint;
}
async rollbackToCheckpoint(
sessionId: string,
checkpointId: string
): Promise<RollbackResult> {
const checkpoints = this.checkpoints.get(sessionId) || [];
const targetIndex = checkpoints.findIndex(c => c.id === checkpointId);
if (targetIndex === -1) {
throw new Error('Checkpoint not found');
}
const checkpointsToRollback = checkpoints.slice(targetIndex + 1).reverse();
const rollbackResults: RollbackStepResult[] = [];
for (const checkpoint of checkpointsToRollback) {
if (!checkpoint.canRollback) {
return {
success: false,
error: `Cannot rollback step: ${checkpoint.stepDescription}`,
partialRollback: rollbackResults
};
}
// Execute rollback actions in reverse order
for (const action of checkpoint.rollbackActions.reverse()) {
try {
await this.executeRollbackAction(sessionId, action);
rollbackResults.push({
checkpointId: checkpoint.id,
action,
success: true
});
} catch (error) {
return {
success: false,
error: `Rollback failed: ${error.message}`,
partialRollback: rollbackResults
};
}
}
}
// Update checkpoint list
this.checkpoints.set(sessionId, checkpoints.slice(0, targetIndex + 1));
return { success: true, rollbackResults };
}
private async captureFileSystemDelta(
sessionId: string,
previousCheckpoint: Checkpoint | null
): Promise<FileSystemDelta> {
const sandbox = await this.sandboxManager.get(sessionId);
const currentState = await sandbox.getFileSystemState();
if (!previousCheckpoint) {
return {
created: [],
modified: [],
deleted: [],
baseState: currentState
};
}
const previousState = previousCheckpoint.fileSystemDelta.baseState;
return {
created: this.findCreatedFiles(previousState, currentState),
modified: this.findModifiedFiles(previousState, currentState),
deleted: this.findDeletedFiles(previousState, currentState),
baseState: currentState
};
}
}
// React hook for checkpoint management
function useCheckpoints(sessionId: string) {
const checkpoints = useAgentStore(
state => state.sessions.get(sessionId)?.checkpoints || []
);
const rollbackToCheckpoint = useCallback(async (checkpointId: string) => {
const result = await checkpointManager.rollbackToCheckpoint(
sessionId,
checkpointId
);
if (result.success) {
// Update agent state to match checkpoint
useAgentStore.getState().syncToCheckpoint(sessionId, checkpointId);
}
return result;
}, [sessionId]);
return {
checkpoints,
rollbackToCheckpoint,
canRollbackTo: (checkpointId: string) => {
const index = checkpoints.findIndex(c => c.id === checkpointId);
return checkpoints.slice(index + 1).every(c => c.canRollback);
}
};
}
Agent Memory Architecture
GIF via GIPHY
interface AgentMemory {
// Short-term: Current task context
shortTerm: {
currentGoal: string;
currentPlan: AgentPlan;
recentActions: ExecutionResult[];
workingContext: Record<string, unknown>;
};
// Long-term: Persistent knowledge
longTerm: {
learnedPatterns: Pattern[];
userPreferences: Preference[];
projectContext: ProjectContext;
pastMistakes: Mistake[];
};
// Episodic: Session-specific
episodic: {
sessionHistory: SessionEvent[];
recoveredErrors: RecoveredError[];
userFeedback: Feedback[];
};
}
class AgentMemoryManager {
private shortTermCache = new Map<string, ShortTermMemory>();
private longTermStore: VectorStore;
private episodicStore: TimescaleDB;
async buildContextForPlanning(
sessionId: string,
goal: string
): Promise<PlanningContext> {
const shortTerm = this.shortTermCache.get(sessionId);
// Retrieve relevant long-term memories
const goalEmbedding = await this.embedder.embed(goal);
const relevantMemories = await this.longTermStore.search(goalEmbedding, {
limit: 10,
filter: { type: 'learned_pattern' }
});
// Get recent episodic context
const recentEpisodes = await this.episodicStore.query({
sessionId,
limit: 20,
orderBy: 'timestamp DESC'
});
// Retrieve past mistakes for similar goals
const pastMistakes = await this.longTermStore.search(goalEmbedding, {
limit: 5,
filter: { type: 'mistake' }
});
return {
goal,
shortTermContext: shortTerm?.workingContext || {},
relevantPatterns: relevantMemories.map(m => m.payload),
recentHistory: recentEpisodes,
mistakesToAvoid: pastMistakes.map(m => m.payload),
userPreferences: await this.getUserPreferences(sessionId)
};
}
async learnFromExecution(
sessionId: string,
execution: ExecutionResult
): Promise<void> {
// Extract learnable patterns
if (execution.status === 'success') {
const pattern = await this.extractPattern(execution);
if (pattern.confidence > 0.7) {
await this.longTermStore.insert({
type: 'learned_pattern',
embedding: await this.embedder.embed(pattern.description),
payload: pattern
});
}
}
// Record mistakes
if (execution.status === 'failed') {
const mistake = await this.analyzeMistake(execution);
await this.longTermStore.insert({
type: 'mistake',
embedding: await this.embedder.embed(mistake.context),
payload: mistake
});
}
// Update episodic memory
await this.episodicStore.insert({
sessionId,
timestamp: Date.now(),
type: 'execution',
payload: execution
});
}
}
UI Components for Agentic Systems
Workflow Visualization
// Real-time workflow visualization component
function WorkflowViewer({ sessionId }: { sessionId: string }) {
const agent = useAgentStore(state => state.sessions.get(sessionId));
const [expandedSteps, setExpandedSteps] = useState<Set<string>>(new Set());
if (!agent || !agent.plan) {
return <WorkflowSkeleton />;
}
return (
<div className="workflow-viewer">
{/* Goal Header */}
<div className="workflow-goal">
<Target className="icon" />
<span className="goal-text">{agent.goal}</span>
<StatusBadge status={agent.status} />
</div>
{/* Plan Confidence */}
<ConfidenceIndicator
confidence={agent.plan.confidence}
estimatedDuration={agent.plan.estimatedDuration}
/>
{/* Step Timeline */}
<div className="step-timeline">
{agent.plan.steps.map((step, index) => {
const execution = agent.executionHistory.find(
e => e.stepId === step.id
);
const isCurrent = index === agent.currentStep;
const isExpanded = expandedSteps.has(step.id);
return (
<WorkflowStep
key={step.id}
step={step}
execution={execution}
isCurrent={isCurrent}
isExpanded={isExpanded}
onToggleExpand={() => {
setExpandedSteps(prev => {
const next = new Set(prev);
next.has(step.id) ? next.delete(step.id) : next.add(step.id);
return next;
});
}}
onRequestApproval={
step.requiresApproval && isCurrent && agent.status === 'waiting'
? () => approveStep(sessionId, step.id)
: undefined
}
/>
);
})}
</div>
{/* Control Panel */}
<WorkflowControls
agent={agent}
onPause={() => agent.pause()}
onResume={() => agent.resume()}
onCancel={() => agent.cancel()}
/>
</div>
);
}
function WorkflowStep({
step,
execution,
isCurrent,
isExpanded,
onToggleExpand,
onRequestApproval
}: WorkflowStepProps) {
const stepStatus = getStepStatus(step, execution, isCurrent);
return (
<div className={`workflow-step ${stepStatus}`}>
<div className="step-header" onClick={onToggleExpand}>
<StepStatusIcon status={stepStatus} />
<span className="step-description">{step.description}</span>
<RiskBadge level={step.riskLevel} />
{isCurrent && <CurrentStepIndicator />}
<ChevronIcon expanded={isExpanded} />
</div>
{isExpanded && (
<div className="step-details">
{/* Tool Info */}
<div className="tool-info">
<span className="tool-name">{step.tool}</span>
<ParameterViewer parameters={step.parameters} />
</div>
{/* Execution Result */}
{execution && (
<ExecutionResultViewer
result={execution}
onRetry={() => retryStep(step.id)}
/>
)}
{/* Approval UI */}
{onRequestApproval && (
<ApprovalPanel
step={step}
onApprove={onRequestApproval}
onReject={() => rejectStep(step.id)}
onModify={() => openStepEditor(step.id)}
/>
)}
</div>
)}
</div>
);
}
Agent Thought Process Viewer
// Shows the agent's "thinking" in real-time
function AgentThoughts({ sessionId }: { sessionId: string }) {
const thoughts = useAgentStore(
state => state.sessions.get(sessionId)?.thoughts || []
);
const isThinking = useAgentStore(
state => state.sessions.get(sessionId)?.status === 'planning'
);
const containerRef = useRef<HTMLDivElement>(null);
// Auto-scroll to latest thought
useEffect(() => {
if (containerRef.current) {
containerRef.current.scrollTop = containerRef.current.scrollHeight;
}
}, [thoughts.length]);
return (
<div className="agent-thoughts" ref={containerRef}>
<div className="thoughts-header">
<Brain className="icon" />
<span>Agent Reasoning</span>
{isThinking && <ThinkingIndicator />}
</div>
<div className="thoughts-list">
{thoughts.map((thought, index) => (
<ThoughtBubble
key={index}
thought={thought}
isLatest={index === thoughts.length - 1}
/>
))}
</div>
</div>
);
}
interface Thought {
type: 'observation' | 'reasoning' | 'decision' | 'action' | 'reflection';
content: string;
timestamp: number;
metadata?: {
confidence?: number;
alternatives?: string[];
reasoning?: string;
};
}
function ThoughtBubble({ thought, isLatest }: ThoughtBubbleProps) {
const icons = {
observation: Eye,
reasoning: Lightbulb,
decision: GitBranch,
action: Play,
reflection: MessageCircle
};
const Icon = icons[thought.type];
return (
<div className={`thought-bubble ${thought.type} ${isLatest ? 'latest' : ''}`}>
<Icon className="thought-icon" />
<div className="thought-content">
<span className="thought-text">{thought.content}</span>
{thought.metadata?.confidence && (
<span className="confidence">
{Math.round(thought.metadata.confidence * 100)}% confident
</span>
)}
{thought.metadata?.alternatives && (
<AlternativesDropdown alternatives={thought.metadata.alternatives} />
)}
</div>
<span className="thought-time">
{formatRelativeTime(thought.timestamp)}
</span>
</div>
);
}
GIF via GIPHY
Checkpoint Timeline
function CheckpointTimeline({ sessionId }: { sessionId: string }) {
const { checkpoints, rollbackToCheckpoint, canRollbackTo } = useCheckpoints(sessionId);
const [selectedCheckpoint, setSelectedCheckpoint] = useState<string | null>(null);
const [isRollingBack, setIsRollingBack] = useState(false);
const handleRollback = async (checkpointId: string) => {
if (!canRollbackTo(checkpointId)) {
toast.error('Cannot rollback to this checkpoint - some steps are irreversible');
return;
}
setIsRollingBack(true);
try {
const result = await rollbackToCheckpoint(checkpointId);
if (result.success) {
toast.success('Rolled back successfully');
} else {
toast.error(result.error);
}
} finally {
setIsRollingBack(false);
}
};
return (
<div className="checkpoint-timeline">
<div className="timeline-header">
<History className="icon" />
<span>Checkpoints</span>
</div>
<div className="timeline-track">
{checkpoints.map((checkpoint, index) => (
<div
key={checkpoint.id}
className={`checkpoint-node ${
canRollbackTo(checkpoint.id) ? 'rollbackable' : 'locked'
}`}
onClick={() => setSelectedCheckpoint(checkpoint.id)}
>
<div className="checkpoint-dot" />
<div className="checkpoint-info">
<span className="checkpoint-step">{checkpoint.stepDescription}</span>
<span className="checkpoint-time">
{formatTime(checkpoint.timestamp)}
</span>
</div>
{!checkpoint.canRollback && (
<Lock className="lock-icon" title="Irreversible step" />
)}
</div>
))}
</div>
{/* Rollback confirmation modal */}
{selectedCheckpoint && (
<RollbackConfirmModal
checkpoint={checkpoints.find(c => c.id === selectedCheckpoint)!}
stepsToRollback={checkpoints.filter(
c => c.timestamp > checkpoints.find(
cp => cp.id === selectedCheckpoint
)!.timestamp
)}
isLoading={isRollingBack}
onConfirm={() => handleRollback(selectedCheckpoint)}
onCancel={() => setSelectedCheckpoint(null)}
/>
)}
</div>
);
}
Safety & Control Architecture
Permission System
interface Permission {
resource: string;
action: 'read' | 'write' | 'execute' | 'delete';
scope: 'global' | 'session' | 'step';
conditions?: PermissionCondition[];
}
interface PermissionCondition {
type: 'path_pattern' | 'rate_limit' | 'time_window' | 'requires_approval';
config: Record<string, unknown>;
}
class PermissionManager {
private userPermissions = new Map<string, Permission[]>();
private sessionOverrides = new Map<string, Permission[]>();
private approvalQueue = new Map<string, ApprovalRequest>();
async check(
userId: string,
tool: string,
parameters: Record<string, unknown>,
sessionId?: string
): Promise<PermissionCheckResult> {
const toolConfig = this.toolRegistry.get(tool);
const requiredPermissions = toolConfig.permissions;
const userPerms = this.userPermissions.get(userId) || [];
const sessionPerms = sessionId
? this.sessionOverrides.get(sessionId) || []
: [];
const allPerms = [...userPerms, ...sessionPerms];
for (const required of requiredPermissions) {
const hasPermission = this.matchPermission(required, allPerms, parameters);
if (!hasPermission) {
return {
allowed: false,
reason: `Missing permission: ${required}`,
canRequest: true
};
}
// Check conditions
const conditions = this.getConditions(required, allPerms);
for (const condition of conditions) {
const conditionResult = await this.evaluateCondition(
condition,
userId,
tool,
parameters
);
if (!conditionResult.passed) {
if (condition.type === 'requires_approval') {
return {
allowed: false,
reason: 'Requires approval',
requiresApproval: true,
approvalContext: conditionResult.context
};
}
return {
allowed: false,
reason: conditionResult.reason
};
}
}
}
return { allowed: true };
}
async requestApproval(
userId: string,
sessionId: string,
tool: string,
parameters: Record<string, unknown>,
context: ApprovalContext
): Promise<ApprovalRequest> {
const request: ApprovalRequest = {
id: generateId(),
userId,
sessionId,
tool,
parameters,
context,
status: 'pending',
createdAt: Date.now(),
expiresAt: Date.now() + 300000 // 5 minute expiry
};
this.approvalQueue.set(request.id, request);
this.notifyApprovers(request);
return request;
}
async approveRequest(
requestId: string,
approverId: string,
modifications?: Record<string, unknown>
): Promise<void> {
const request = this.approvalQueue.get(requestId);
if (!request) throw new Error('Request not found');
if (request.status !== 'pending') throw new Error('Request already processed');
request.status = 'approved';
request.approverId = approverId;
request.approvedAt = Date.now();
if (modifications) {
request.modifiedParameters = modifications;
}
// Grant temporary session permission
this.grantSessionPermission(request.sessionId, {
resource: request.tool,
action: 'execute',
scope: 'step',
conditions: [{
type: 'rate_limit',
config: { max: 1 }
}]
});
this.emitApprovalResult(request);
}
}
Sandbox Execution
GIF via GIPHY
interface SandboxConfig {
type: 'container' | 'vm' | 'wasm' | 'process';
resources: {
cpuLimit: number; // Percentage
memoryLimit: number; // MB
diskLimit: number; // MB
networkAccess: boolean;
allowedHosts?: string[];
};
filesystem: {
rootPath: string;
writablePaths: string[];
readOnlyPaths: string[];
blockedPaths: string[];
};
timeout: number;
}
class SandboxManager {
private activeSandboxes = new Map<string, Sandbox>();
async createSandbox(sessionId: string, config: SandboxConfig): Promise<Sandbox> {
let sandbox: Sandbox;
switch (config.type) {
case 'container':
sandbox = await this.createContainerSandbox(config);
break;
case 'wasm':
sandbox = await this.createWasmSandbox(config);
break;
case 'process':
sandbox = await this.createProcessSandbox(config);
break;
default:
throw new Error(`Unknown sandbox type: ${config.type}`);
}
this.activeSandboxes.set(sessionId, sandbox);
// Set up resource monitoring
this.monitorResources(sessionId, sandbox, config);
return sandbox;
}
private async createContainerSandbox(config: SandboxConfig): Promise<Sandbox> {
const container = await docker.createContainer({
Image: 'agent-sandbox:latest',
HostConfig: {
Memory: config.resources.memoryLimit * 1024 * 1024,
CpuPeriod: 100000,
CpuQuota: config.resources.cpuLimit * 1000,
NetworkMode: config.resources.networkAccess ? 'bridge' : 'none',
Binds: this.buildMounts(config.filesystem),
ReadonlyRootfs: true,
SecurityOpt: ['no-new-privileges']
}
});
await container.start();
return new ContainerSandbox(container, config);
}
private monitorResources(
sessionId: string,
sandbox: Sandbox,
config: SandboxConfig
): void {
const interval = setInterval(async () => {
const stats = await sandbox.getStats();
if (stats.memoryUsage > config.resources.memoryLimit * 0.9) {
this.emit('resource_warning', {
sessionId,
type: 'memory',
usage: stats.memoryUsage,
limit: config.resources.memoryLimit
});
}
if (stats.cpuUsage > config.resources.cpuLimit * 0.9) {
this.emit('resource_warning', {
sessionId,
type: 'cpu',
usage: stats.cpuUsage,
limit: config.resources.cpuLimit
});
}
}, 1000);
sandbox.on('exit', () => clearInterval(interval));
}
}
class ContainerSandbox implements Sandbox {
constructor(
private container: Container,
private config: SandboxConfig
) {}
async exec(
command: string,
options: ExecOptions = {}
): Promise<ExecResult> {
const exec = await this.container.exec({
Cmd: ['sh', '-c', command],
WorkingDir: options.cwd || '/workspace',
AttachStdout: true,
AttachStderr: true
});
const startTime = Date.now();
const timeout = options.timeout || this.config.timeout;
return new Promise((resolve, reject) => {
const timer = setTimeout(() => {
exec.kill('SIGKILL');
reject(new Error('Execution timeout'));
}, timeout);
exec.start({ hijack: true }, (err, stream) => {
if (err) {
clearTimeout(timer);
reject(err);
return;
}
let stdout = '';
let stderr = '';
stream.on('data', (chunk: Buffer) => {
// Docker multiplexes stdout/stderr
const type = chunk[0];
const payload = chunk.slice(8).toString();
if (type === 1) stdout += payload;
else if (type === 2) stderr += payload;
// Check buffer limits
if (stdout.length + stderr.length > (options.maxBuffer || 1048576)) {
exec.kill('SIGKILL');
reject(new Error('Output buffer exceeded'));
}
});
stream.on('end', async () => {
clearTimeout(timer);
const inspect = await exec.inspect();
resolve({
stdout,
stderr,
exitCode: inspect.ExitCode,
duration: Date.now() - startTime
});
});
});
});
}
async readFile(path: string): Promise<string> {
this.validatePath(path, 'read');
const result = await this.exec(`cat "${path}"`);
if (result.exitCode !== 0) {
throw new Error(`Failed to read file: ${result.stderr}`);
}
return result.stdout;
}
async writeFile(path: string, content: string): Promise<void> {
this.validatePath(path, 'write');
// Use heredoc for safe content handling
const escapedContent = content.replace(/'/g, "'\\''");
const result = await this.exec(`cat > "${path}" << 'SANDBOX_EOF'
${content}
SANDBOX_EOF`);
if (result.exitCode !== 0) {
throw new Error(`Failed to write file: ${result.stderr}`);
}
}
private validatePath(path: string, operation: 'read' | 'write'): void {
const normalizedPath = nodePath.normalize(path);
// Check blocked paths
for (const blocked of this.config.filesystem.blockedPaths) {
if (normalizedPath.startsWith(blocked)) {
throw new Error(`Access denied: ${path}`);
}
}
// Check write permissions
if (operation === 'write') {
const isWritable = this.config.filesystem.writablePaths.some(
wp => normalizedPath.startsWith(wp)
);
if (!isWritable) {
throw new Error(`Write access denied: ${path}`);
}
}
}
}
Observability & Monitoring
Agent Telemetry
interface AgentTelemetry {
// Session metrics
sessionsActive: Gauge;
sessionsTotal: Counter;
sessionDuration: Histogram;
// Planning metrics
planGenerationTime: Histogram;
planStepCount: Histogram;
replanCount: Counter;
// Execution metrics
stepExecutionTime: Histogram;
stepSuccessRate: Counter;
toolCallsPerSession: Histogram;
// Safety metrics
permissionDenials: Counter;
approvalRequests: Counter;
rollbacksPerformed: Counter;
// Resource metrics
sandboxCpuUsage: Gauge;
sandboxMemoryUsage: Gauge;
checkpointSize: Histogram;
}
class AgentTelemetryCollector {
private metrics: AgentTelemetry;
private tracer: Tracer;
async traceAgentSession(
sessionId: string,
agentLoop: AsyncIterable<AgentEvent>
): Promise<void> {
const sessionSpan = this.tracer.startSpan('agent.session', {
attributes: { sessionId }
});
const sessionStart = Date.now();
this.metrics.sessionsActive.add(1);
this.metrics.sessionsTotal.add(1);
let stepCount = 0;
let replanCount = 0;
let toolCallCount = 0;
try {
for await (const event of agentLoop) {
switch (event.type) {
case 'planning_started':
sessionSpan.addEvent('planning_started');
break;
case 'plan_generated':
this.metrics.planGenerationTime.record(event.duration);
this.metrics.planStepCount.record(event.plan.steps.length);
sessionSpan.addEvent('plan_generated', {
stepCount: event.plan.steps.length,
confidence: event.plan.confidence
});
break;
case 'step_started':
stepCount++;
toolCallCount++;
sessionSpan.addEvent('step_started', { stepId: event.step.id });
break;
case 'step_completed':
this.metrics.stepExecutionTime.record(event.duration);
this.metrics.stepSuccessRate.add(1, {
status: event.result.status
});
break;
case 'replanned':
replanCount++;
this.metrics.replanCount.add(1);
break;
case 'permission_denied':
this.metrics.permissionDenials.add(1, {
tool: event.tool,
reason: event.reason
});
break;
case 'rollback':
this.metrics.rollbacksPerformed.add(1, {
steps: event.stepsRolledBack
});
break;
}
}
sessionSpan.setStatus({ code: SpanStatusCode.OK });
} catch (error) {
sessionSpan.setStatus({
code: SpanStatusCode.ERROR,
message: error.message
});
throw error;
} finally {
this.metrics.sessionsActive.add(-1);
this.metrics.sessionDuration.record(Date.now() - sessionStart);
this.metrics.toolCallsPerSession.record(toolCallCount);
sessionSpan.setAttribute('stepCount', stepCount);
sessionSpan.setAttribute('replanCount', replanCount);
sessionSpan.setAttribute('toolCallCount', toolCallCount);
sessionSpan.end();
}
}
}
// Production dashboard configuration
const agentDashboard = {
panels: [
{
title: 'Agent Activity',
metrics: [
'agent_sessions_active',
'agent_sessions_total_rate',
'agent_session_duration_p50',
'agent_session_duration_p99'
]
},
{
title: 'Execution Health',
metrics: [
'agent_step_success_rate',
'agent_step_execution_time_p50',
'agent_replan_rate',
'agent_rollbacks_rate'
]
},
{
title: 'Safety Metrics',
metrics: [
'agent_permission_denials_rate',
'agent_approval_requests_rate',
'agent_sandbox_violations',
'agent_timeout_rate'
]
},
{
title: 'Resource Usage',
metrics: [
'agent_sandbox_cpu_avg',
'agent_sandbox_memory_avg',
'agent_checkpoint_size_avg',
'agent_context_tokens_avg'
]
}
],
alerts: [
{
name: 'High Failure Rate',
condition: 'agent_step_success_rate < 0.8',
severity: 'critical'
},
{
name: 'Permission Denials Spike',
condition: 'rate(agent_permission_denials) > 10',
severity: 'warning'
},
{
name: 'Sandbox Resource Exhaustion',
condition: 'agent_sandbox_memory_avg > 0.9',
severity: 'critical'
}
]
};
GIF via GIPHY
Production Incidents & Lessons
Incident 1: Runaway Agent Loop
Symptoms: Agent entered infinite replan loop, consuming significant LLM tokens and never completing.
Root Cause: The agent's success criteria were ambiguous. Each replan slightly modified the approach, but the goal validation always returned "not quite complete."
// Before: Vague completion check
async function isGoalComplete(agent: AgentLoop): Promise<boolean> {
const response = await llm.complete({
messages: [{
role: 'user',
content: `Is this goal complete? Goal: ${agent.goal}\nCurrent state: ${agent.currentState}`
}]
});
return response.content.toLowerCase().includes('yes');
}
// After: Structured completion criteria
interface GoalCompletion {
criteria: CompletionCriterion[];
requiredCriteria: number; // e.g., 3 out of 4 must pass
maxAttempts: number;
}
async function isGoalComplete(
agent: AgentLoop,
completion: GoalCompletion
): Promise<CompletionResult> {
const results = await Promise.all(
completion.criteria.map(c => evaluateCriterion(c, agent))
);
const passedCount = results.filter(r => r.passed).length;
// Hard stop after max attempts
if (agent.replanCount >= completion.maxAttempts) {
return {
complete: true,
reason: 'max_attempts_reached',
partialSuccess: passedCount > 0
};
}
return {
complete: passedCount >= completion.requiredCriteria,
passedCriteria: results.filter(r => r.passed),
failedCriteria: results.filter(r => !r.passed)
};
}
Prevention: Always define explicit, measurable completion criteria. Implement hard limits on replans and total execution time.
Incident 2: Checkpoint Storage Exhaustion
Symptoms: Agent sessions started failing with "storage quota exceeded" errors. Investigation revealed checkpoint storage grew to 500GB.
GIF via GIPHY
Root Cause: Each file write created a full file backup in the checkpoint. Large file operations (downloading dependencies, building artifacts) created massive checkpoints.
// Before: Full file backup
async function createFileBackup(path: string): Promise<Backup> {
const content = await fs.readFile(path);
return { path, content, size: content.length };
}
// After: Delta-based checkpoints with size limits
class DeltaCheckpointManager {
private maxCheckpointSize = 10 * 1024 * 1024; // 10MB
async createCheckpoint(
sessionId: string,
changes: FileChange[]
): Promise<Checkpoint> {
const deltas: FileDelta[] = [];
let totalSize = 0;
for (const change of changes) {
// Skip large files - mark as non-rollbackable
if (change.size > this.maxCheckpointSize / 10) {
deltas.push({
path: change.path,
type: 'large_file',
rollbackable: false
});
continue;
}
// Use diff for text files
if (this.isTextFile(change.path)) {
const diff = this.computeDiff(change.oldContent, change.newContent);
if (diff.length < change.newContent.length * 0.5) {
deltas.push({
path: change.path,
type: 'diff',
diff,
rollbackable: true
});
totalSize += diff.length;
continue;
}
}
// Full backup for small binary files
deltas.push({
path: change.path,
type: 'full',
content: change.oldContent,
rollbackable: true
});
totalSize += change.oldContent?.length || 0;
// Stop if checkpoint too large
if (totalSize > this.maxCheckpointSize) {
break;
}
}
return { deltas, totalSize, partialCheckpoint: totalSize > this.maxCheckpointSize };
}
}
Incident 3: Tool Injection Attack
Symptoms: Security audit revealed agent could be tricked into executing arbitrary commands by manipulating its context.
Root Cause: User-provided input in the goal was directly interpolated into tool parameters without sanitization.
// Before: Direct interpolation
const goal = userInput; // "Create a file named `; rm -rf /; echo `test.txt"
const step = {
tool: 'write_file',
parameters: { path: `./output/${goal.split(' ').pop()}` }
};
// After: Strict parameter validation and sandboxing
class SecureToolExecutor {
async execute(
tool: Tool,
parameters: Record<string, unknown>,
context: ExecutionContext
): Promise<ToolResult> {
// 1. Schema validation
const validationResult = this.validateParameters(tool, parameters);
if (!validationResult.valid) {
throw new ValidationError(validationResult.errors);
}
// 2. Sanitization
const sanitizedParams = this.sanitizeParameters(tool, parameters);
// 3. Permission check
const permissionResult = await this.permissionManager.check(
context.userId,
tool.name,
sanitizedParams
);
if (!permissionResult.allowed) {
throw new PermissionError(permissionResult.reason);
}
// 4. Execute in sandbox with minimal privileges
const result = await this.sandbox.execute(() =>
tool.execute(sanitizedParams, context)
);
// 5. Validate output
this.validateOutput(tool, result);
return result;
}
private sanitizeParameters(
tool: Tool,
parameters: Record<string, unknown>
): Record<string, unknown> {
const sanitized: Record<string, unknown> = {};
for (const [key, value] of Object.entries(parameters)) {
const schema = tool.parameters.properties[key];
if (schema.type === 'string') {
// Remove shell metacharacters
sanitized[key] = this.sanitizeString(value as string);
} else {
sanitized[key] = value;
}
}
return sanitized;
}
private sanitizeString(value: string): string {
// Remove dangerous patterns
return value
.replace(/[;&|`$(){}[\]<>]/g, '')
.replace(/\.\./g, '')
.slice(0, 1000); // Length limit
}
}
Tradeoffs & Engineering Decisions
Decision: Synchronous vs Asynchronous Tool Execution
| Factor | Synchronous | Asynchronous |
|---|---|---|
| User feedback | Immediate | Delayed |
| Complexity | Simple | Complex |
| Parallel execution | Not possible | Possible |
| Error handling | Straightforward | Complex (partial failures) |
| Rollback | Simpler | Complex (dependency graph) |
Decision: Hybrid approach. Default to synchronous with async opt-in for independent operations. The agent explicitly declares parallel-safe steps.
Decision: Client-side vs Server-side Agent Loop
Client-side:
- Pros: Real-time UI updates, offline capability, reduced server load
- Cons: Security risks, limited resource access, browser constraints
Server-side:
GIF via GIPHY
- Pros: Secure execution, full resource access, consistent environment
- Cons: Latency for UI updates, server scaling, state synchronization
Decision: Server-side agent loop with real-time event streaming to client. The client only handles UI rendering and user input—all planning, execution, and state management happens server-side.
Decision: Granularity of Checkpoints
| Approach | Pros | Cons |
|---|---|---|
| Per-step | Fine-grained rollback | Storage overhead |
| Per-plan | Less storage | Coarse rollback |
| On-demand | Minimal storage | May miss important states |
Decision: Per-step checkpoints for reversible operations, with delta compression. Mark irreversible operations explicitly and warn users before execution.
Future Evolution
Multi-Agent Coordination
interface AgentTeam {
coordinator: Agent;
specialists: Agent[];
sharedContext: SharedContext;
communicationChannel: MessageChannel;
}
// Example: Code review team
const codeReviewTeam: AgentTeam = {
coordinator: { role: 'lead_reviewer', capabilities: ['planning', 'synthesis'] },
specialists: [
{ role: 'security_reviewer', capabilities: ['security_analysis'] },
{ role: 'performance_reviewer', capabilities: ['perf_analysis'] },
{ role: 'style_reviewer', capabilities: ['lint', 'formatting'] }
],
sharedContext: new SharedContext(),
communicationChannel: new BroadcastChannel('review-team')
};
Human-in-the-Loop Learning
interface FeedbackLoop {
captureUserCorrection(
stepId: string,
originalAction: Action,
correctedAction: Action
): Promise<void>;
learnFromFeedback(): Promise<void>;
applyLearnings(context: PlanningContext): Promise<PlanModification[]>;
}
GIF via GIPHY
Proactive Agents
Moving from reactive (user-initiated) to proactive (agent-initiated) workflows:
interface ProactiveAgent {
// Monitor triggers
triggers: Trigger[];
// Autonomous actions within boundaries
autonomyBounds: AutonomyConfig;
// Self-initiated workflows
initiateWorkflow(trigger: TriggerEvent): Promise<AgentLoop>;
}
Conclusion
Building agentic UI systems requires rethinking fundamental frontend assumptions. When the interface hosts an autonomous agent that plans, executes, and adapts, traditional patterns for state management, error handling, and user interaction become insufficient.
Key architectural principles:
GIF via GIPHY
- Safety by design: Sandboxing, permissions, and approval gates aren't afterthoughts—they're core architecture
- Observable autonomy: Users must understand what the agent is doing, why, and how to intervene
- Graceful rollback: Every action should be reversible where possible, with clear indication when it's not
- Structured completion: Define explicit success criteria to prevent runaway execution
- Defense in depth: Multiple layers of validation, sanitization, and containment
The future of agentic UI involves multi-agent teams, proactive workflows, and increasingly sophisticated human-AI collaboration patterns. The architectural foundations established today—around safety, observability, and user control—will determine how effectively these more advanced systems can be deployed in production.
What did you think?