SXF RESEARCH / CODING AGENT SYSTEMS

AI Coding Agent Architecture:
Tools vs Subagents, Plugins, Functions & Actions

What really happens between “fix my repository” and a verified pull request? A technical reference to the control loop behind modern AI coding agents: tools, function calls, plugins, skills, subagents, MCP, permissions, sandboxes, evaluation and proof of execution.

SXF architectural illustration showing intent, agent harness, delegated subagents, callable tools, plugins and governed code execution
The AI coding agent is a controlled execution system—not a language model with extra buttons. Diagram: SXF original conceptual model, not a vendor benchmark.
QUICK ANSWER

Tools execute capabilities; subagents own delegated tasks; functions expose host handlers; plugins package extensions; actions produce effects.

These concepts describe different architectural layers. Skills encode repeatable instructions, MCP standardizes access to capabilities, and the agent harness manages context, selection, recovery and verification. Crucially, a model's tool request is not itself authority to act.

DEFINITION

What is AI coding agent architecture?

AI coding agent architecture is the design of the control loop, context pipeline, execution environment, capability interfaces, delegated workers, authority boundaries and verification process that turn a software-development objective into a reviewable, testable change. The language model is an important component, but it is not the architecture.

A basic coding assistant predicts or edits code in response to a user. A coding agent can inspect a repository, select tools, execute commands, respond to failing tests, change state and continue until it reaches a stopping condition. This makes the critical questions operational: What can it observe? What may it change? Who authorizes those changes? How does the system prove its work?

CONTROL PLANE

The decision loop

The model and harness inspect context, plan, select capabilities, receive tool results and decide whether to continue.

EXECUTION PLANE

The real environment

A sandbox or application runtime reads files, runs commands, invokes APIs and produces observable effects.

AUTHORITY PLANE

Who can do what?

Identity, permission scopes, tool allowlists, approvals and audit logs restrict actions outside the model's discretion.

EVIDENCE PLANE

What actually succeeded?

Tests, diffs, exit codes, logs, artifacts and human review substantiate outcomes rather than trusting claims.

SXF architecture principleThe model can propose an operation. The authorized host, service or tool performs it. Model-generated text never creates permission by itself.

SEARCH-INTENT COMPARISON

AI coding agent platform concepts: tools vs subagents vs plugins vs functions vs actions

The key terms appear together in product menus and documentation, but they represent different layers of the system. An apples-to-apples comparison begins by asking whether an item is a worker, a capability, an invocation interface, a distribution mechanism or an effect.

ConceptTechnical roleWhat it ownsCoding exampleDo not confuse with
ToolCallable capabilityOne bounded operation and a resultSearch the repo; run a test; edit a fileAn independent reasoning worker
FunctionApplication-defined call interfaceSchema + host handlerrun_tests(target)Every possible tool type
SubagentDelegated agent executionA scoped subtask, context and multiple stepsRead-only security reviewerA single function/API request
PluginPlatform-defined extension packageDistribution and configurationBundle an MCP connector and a skillA universal protocol primitive
SkillReusable task instructionsProcedural guidance and optional assetsA documented migration workflowAuthorization or tool execution
ActionAttempted state transitionAn operation with side effects and outcomeApply patch; open pull requestInherent proof of authorization
MCP serverInteroperable capability providerDiscoverable tools, context and service-side checksGitHub repository access connectorAn autonomous subagent

These are precise analytical distinctions, not a claim that every vendor uses identical vocabulary. For example, GitHub distinguishes tools, MCP servers, skills, hooks, subagents and custom agents; OpenAI describes plugins as packages that can contain skills and MCP configuration. The authoritative names for a particular product should always come from that product's documentation. [1][2][8][9]

END-TO-END SYSTEM

How does an AI coding agent work? The complete execution loop

Suppose the task is: “Upgrade the authentication dependency, keep existing behavior, and prepare a pull request.” A serious agent architecture must connect intent to executed evidence without silently enlarging the task's scope.

01 · INTAKE

Turn the request into a contract

Resolve repository, base revision, acceptance criteria, allowed file scope and whether edits or external writes are authorized. Stop or ask when the goal is not sufficiently defined.

02 · DISCOVER

Retrieve relevant source evidence

Inspect manifests, lockfiles, repository instructions, call sites, tests and git state. Treat retrieved files, web pages and tool output as untrusted data, not fresh higher-priority instructions.

03 · PLAN

Make steps measurable

Identify dependent migrations and validation work, and decide whether specialist research is separable enough to delegate. A plan must include a stopping condition.

04 · EXECUTE

Run approved operations

Call read, edit and test tools through a constrained executor. Capture arguments, principal, environment, output and side effects. Timeouts and permission denials are normal results.

05 · VERIFY

Prove the implementation

Run targeted tests, inspect changes, check lockfiles and identify untested paths. A generated summary is not evidence that commands were actually run.

06 · DELIVER

Make review possible

Create a patch or pull request with commands executed, observed pass/fail results, changed files and outstanding risks. Keep merge and deploy permissions separately gated.

OpenAI's managed-agent documentation explicitly separates the harness controlling the loop, an optional environment for files and execution, and the application server integrating external function tools and the product. This decomposition is useful even when another vendor ships all three together. [5][6][7]

CAPABILITY INTERFACES

AI agent tools vs function calling: what actually executes?

A tool is an interface the agent may invoke to read information or effect a bounded operation. The implementation could be internal to the harness (terminal, patch, browser), provided as a host-application function, or exposed by an MCP server. Function calling is a common structured invocation mechanism: the model proposes a function name and arguments, while host code validates the request, runs the handler and supplies its result.

The language model does not execute arbitrary operations merely by emitting a JSON object. A model can request run_tests; the executor must validate the path, ensure the caller is allowed, run the process under limits and return actual exit status.

# Illustrative pseudocode: not a vendor SDK signature
def handle_tool_call(request, identity, sandbox):
    target = validate_test_target(request.arguments["target"])
    policy.require(identity, "tests.execute", target)
    result = sandbox.run(["test-runner", target], timeout=120)
    return {
        "exit_code": result.exit_code,
        "report": result.report_artifact,
        "stdout": redact_and_truncate(result.stdout)
    }
# Tool invocation is a request; execution and permission live in the host.

Good interfaces expose typed inputs, stable names, timeouts, output limits, retry rules and explicit side-effect classification. A correct schema catches malformed arguments; it does not determine that an otherwise well-formed operation should be allowed.

Is a shell command a tool or an action?

The shell executor is a tool. A particular invocation such as git diff --check is an action requested through it. The exit status, output and changed repository state are results. Logging these separately is essential for post-run verification.

PROTOCOL LAYER

How MCP fits into AI coding agent architecture

The Model Context Protocol (MCP) is an interoperability boundary that lets a client discover and invoke capabilities exposed by a server. MCP tools have names and schemas, and the client can use tools/list to discover them and tools/call to invoke them. An MCP server can also expose other context capabilities, which are not automatically executable tools. [11]

The 2026-07-28 MCP specification is materially different from older examples that assume legacy initialization and session behavior. Its stateless requests and metadata rules should be followed for new integrations; version pinning and backward compatibility must be explicit. [12]

CLIENT

Select and mediate

Decides which eligible server tools to expose to the model and handles invocation and responses.

SERVER

Validate and serve

Advertises capabilities; independently checks inputs, identity, scopes and service authorization.

SCHEMAS

Constrain shape

Describe valid tool inputs and, where provided, structured outputs. They are not the permission system.

TRUST BOUNDARY

Enforce policy

Treat tool results and annotations carefully; get confirmation for sensitive operations, with audit records.

MCP is neither a subagent nor a plugin by definition. A plugin may configure an MCP server. A subagent may invoke MCP tools. Those are composition choices. See the site's deeper Model Context Protocol guide for transport, OAuth, resources, prompts and server architecture.

DELEGATION STRATEGY

Subagents vs tools: when should you create another coding agent?

A subagent is a delegated worker with a defined objective, a distinct execution context and, ideally, intentionally restricted capabilities. It may perform many tool calls to deliver a review, report, patch or research result. A tool call, by contrast, is a bounded capability invocation. GitHub's Copilot SDK documentation illustrates specialist agents with independent prompts and tool restrictions, and explicitly describes subagent lifecycle events. [4]

KEEP ONE AGENT

Low coordination cost

Prefer one agent for small sequential fixes, tightly coupled file edits and work that already fits a clear verification loop.

DELEGATE RESEARCH

Isolated exploration

Use read-only workers to map APIs, inspect tests or research migrations without sharing write permissions.

DELEGATE REVIEW

Independent check

Give a security or test reviewer a fixed diff and rubric. Do not assume independence if it inherits the writer's unsupported conclusions.

LIMIT PARALLEL EDITS

Avoid merge hazards

Give each writer isolated branches or worktrees and explicit file ownership; validate the integrated patch afterward.

Extra agents cost tokens, initialization, duplicated retrieval, coordination time and conflict resolution. Parallelism is valuable only when subtasks can proceed independently and a final integrator checks the combined output. A specialist without a measurable deliverable is often overhead, not architectural sophistication.

Decision ruleAdd a subagent only when a task has a separable objective, a checkable return contract, a reason for independent context or permissions, and expected value greater than coordination cost.

EXTENSION MODEL

Plugins vs skills vs hooks vs custom agents explained

A plugin generally packages integrations for distribution or installation; the payload is platform-specific. OpenAI documents plugins packaging skills and MCP configuration. GitHub Copilot's extension documentation describes bundles that may include skills, custom agents, hooks and MCP server configurations. Neither makes “plugin” a synonym for a remote reasoning agent. [8][2]

A skill stores task guidance, potentially supported by reference files and scripts. It tells an agent how to approach work; it does not give the agent credentials or bypass policy. A hook is application-controlled code invoked at a lifecycle event rather than a new free-form model-selected agent. A custom agent profile defines specialist instructions and tools; a subagent is a particular delegated run. These distinctions prevent both unsafe configuration and inflated product claims.

Original SXF diagram comparing subagents, tool functions, plugins, skills and actions by their separate architectural jobs
SXF original reference taxonomy: workers, call interfaces, extension packages, procedural instructions and state changes are separate concerns.

Do skills or repository instructions give the agent new permissions?

No. A skill saying “merge this code” does not create a valid merge authorization. Tool registration determines which operations are reachable; identity and policy decide what the current principal may do. A repository file is also untrusted input relative to the application's privileged instructions. Related context: GitHub Copilot Memory discusses durable repository knowledge separately from instructions.

SIDE EFFECTS

What are actions in AI coding agents?

An action is the system's attempted operation: run a compiler, apply a patch, make a network request, delete a branch or open a pull request. The word may also name a vendor feature—GitHub Actions is a specific CI/CD product—so technical writing must distinguish general actions from branded ones.

Record three levels independently: intent (“test package A”), execution (tool, arguments, principal, environment, time and exit code) and outcome (logs, artifacts, test verdict or changed state). An agent claiming that “tests passed” is not evidence that a test command ran successfully. The executor's observed report is evidence.

Actions should also be classified by side effects: read-only, local reversible write, external write and irreversible or production-affecting operations. These classes guide approval gates and retry rules; repeating a failed read is not equivalent to retrying a payment, merge or deployment.

PRODUCTION PATTERN

A practical AI coding agent reference architecture

The following is an SXF recommended architecture pattern, derived from the verified roles and general isolation principles—not a claim that vendors implement one identical design or that it has been independently benchmarked.

USER GOAL + ACCEPTANCE CRITERIA
              |
       SCOPE / IDENTITY GATE
              |
     ORCHESTRATOR + TASK STATE
       /                 \
READ-ONLY SPECIALIST   SCOPED EXECUTOR
  (optional worker)     (typed tools)
       \                 /
           VERIFICATION
      test | diff | logs | review
                  |
        HUMAN APPROVAL / PR
                  |
          SEPARATE MERGE GATE

Start with one orchestrator and a minimal set of discover, edit and test tools. Run writes in a repository-scoped sandbox, keep external network access explicit, and require checks before proposing a reviewable change. Only then add custom skills, MCP connectors, plugins or specialist subagents.

What should a tool execution record contain?

call:
  tool: repo.run_tests
  principal: task-agent
  target: packages/auth
  workspace: isolated-worktree-73
  timeout_seconds: 120
authorization:
  permission: tests.execute
  decision: allow
result:
  exit_code: 1
  status: failed
  artifact: test-report.json
  state_delta: none
  trace_id: run-73-step-04

The goal is to make failure explicit, observable and recoverable. A failed test is useful feedback; silently rewriting it as success is an architecture failure.

ENGINEERING CASES

Three AI coding agent architectures compared in real workflows

PATTERN A · MINIMAL

Bounded bug fix

One agent + repository search + edit + tests. Strong when the task is sequential and localized. Measure regression-test outcomes, smallest valid diff and review time.

PATTERN B · DELEGATED

Large dependency migration

A parent coordinates read-only compatibility and security reviewers while one writer owns integration. Watch duplicated research, stale context and branch conflicts.

PATTERN C · CONNECTED

Cross-system workflow

An agent uses MCP tools for issues, code hosting and internal docs while service-side policies protect external writes. Require explicit scopes and approvals.

There is no universal winner. The correct architecture depends on task coupling, isolation needs, run length, repository size, cost constraints, permission boundaries and the reliability of the tools behind it.

THREAT MODEL

AI coding agent security: where authority must stop

A coding agent routinely consumes attacker-influenced material: repository comments, issue descriptions, external docs, package metadata, web pages and tool output. A malicious instruction embedded in that material can attempt to redirect the agent to leak a secret or invoke a privileged tool. The defense is not merely to ask the model to ignore suspicious prose; the execution system must enforce the boundary.

01 · PERMISSION

Least-privilege scopes

Separate read, local edit, test execution, network and publish privileges. Credentials belong to the service or bounded executor, not arbitrary prompt text.

02 · SANDBOX

Contain effects

Limit filesystem roots, outbound network, CPU, time and mounted credentials; isolate untrusted tasks from production systems.

03 · APPROVAL

Gate expensive actions

Independently approve merges, deployments, deletion, secret access and high-risk external writes.

04 · PROVENANCE

Audit and validate

Capture which principal called which tool, exact arguments, authorization decision, result and resulting state; validate returned content.

MCP's current tools specification requires server-side input validation and access controls, and recommends client confirmations around sensitive operations. Protocol compliance alone is not a complete security model. Read the companion AI Agent Security and Prompt Injection references. [11]

MEASUREMENT

How to benchmark AI coding agent architecture—not just its model

Hold the task set, repository revision, model budget, timeout, permitted tools, environment and reviewer rubric constant across configurations. Repeat runs. Compare the full system's accepted changes—not screenshots of a successful chat response. Evaluate a single-agent baseline before adding one extension at a time.

MetricWhat it testsRecord
Accepted change rateActual success on the business taskAccepted patches / tasks
Repeated-run consistencyWhether success is dependableRun-level variance and strict pass-all runs
Tool fault recoveryResilience to API errors and failed commandsRecovered faults / injected faults
Side-effect disciplineWhether the agent edits beyond scopeExtraneous files and prohibited operations
P50/P95 latencyTail performance and coordination overheadWall-clock distribution
Total task costWhether delegation is economically usefulTokens, compute, retries, human review time
ReviewabilityQuality of delivered evidenceCorrections, diff clarity, test trace completeness
Safety boundary adherenceWhether policy was enforcedUnauthorized attempts and violations

Run ablations: baseline; baseline + skill; baseline + MCP; baseline + read-only reviewer; baseline + parallel specialists. Attribute any improvement to a mechanism, such as better repository retrieval, clearer requirements, stronger test coverage or reduced latency. Do not assume adding agents increases accuracy. For deeper consistency metrics see AI Agent Reliability; for retrieval and persistent state see AI Agent Memory.

DESIGN ERRORS

Seven AI coding agent architecture mistakes to avoid

  1. Calling every integration an “agent.” A remote API capability is not automatically a delegated worker.
  2. Confusing tool-call syntax with authorization. Typed arguments are necessary but do not establish a right to execute.
  3. Treating skills as runtime permissions. Instructions are not credentials or policy grants.
  4. Assuming MCP provides a trust boundary for free. It standardizes interfaces; clients and services still must verify.
  5. Parallel editing without ownership. Shared worktrees and overlapping patches can invalidate each other's assumptions.
  6. Accepting narrative success without run evidence. A claimed passing test is not an executed passing test.
  7. Comparing models while changing the harness and budget. Without controlled conditions, architecture conclusions are not reproducible.

ARCHITECTURAL DECISION

Which AI coding agent architecture should you build first?

Build the smallest system that can produce an accepted patch with evidence.

Begin with one agent, one bounded repository workspace, explicit read/edit/test tools and an independent review gate. Add a skill for repeatable instructions, an MCP connector for external capabilities, a plugin for packaging integrations and a subagent only when an isolated work package has measurable benefits.

For decisions about actual developer products, see Best AI Coding Tools. For larger orchestration, consult How to Build an AI Super Agent. These references answer different search intentions: product choice, workflow design and the underlying architecture described here.

DIRECT ANSWERS

AI coding agent architecture FAQs

What is AI coding agent architecture?

It is the system design that connects a model-driven control loop to context retrieval, tools, execution environments, delegated workers, permissions, verification and final code artifacts.

What is the difference between tools and subagents?

A tool performs a bounded operation through a defined interface. A subagent is a delegated agent execution with its own task and usually a separate context; it may use several tools before returning a result.

Are tools and functions the same in AI agents?

A function call is one interface for invoking an application-implemented operation. Some tools are functions, but tools may also be shell, browser, search, code execution or MCP capabilities.

What is the difference between plugins, skills and MCP?

A plugin is a platform-defined extension package. A skill contains reusable workflow instructions and optional supporting resources. MCP is a protocol for exposing tools and other context capabilities; the three are not interchangeable.

What are actions in a coding agent?

An action is an attempted operation such as editing a file, executing a test or opening a pull request. A tool exposes the interface, a host executes the action, and an authorization layer decides whether it is allowed.

When should I use coding subagents?

Use subagents when there are independent tasks with clear deliverables, separate context requirements or distinct tool scopes. Avoid delegating work when coordination and reconciliation costs dominate.

Does MCP automatically secure an AI coding agent?

No. MCP provides interoperable interfaces; systems still require identity, narrow permission scopes, sandboxing, validation, audit records and approvals for sensitive side effects.

How should AI coding agents be compared?

Hold tasks and environment constant, repeat evaluations, and measure accepted-change rate, test results, run-to-run variability, tool recovery, unnecessary edits, latency, cost, reviewer burden and boundary violations.

RESEARCH METHODOLOGY

What the evidence supports—and what it does not

This SXF article is an independently authored technical synthesis, last verified October 10, 2026, built from official vendor documentation and the published MCP specification. Our layer model, matrices, sample architecture and decision rules are original explanatory frameworks, not results of a controlled performance experiment. Vendor documentation is evidence for that vendor's interface and terminology, not proof of performance superiority. Implementation details, pricing and feature support may change; the permanent URL deliberately contains no year.

No claim is made that subagents always improve coding accuracy, that one agent framework wins every workload, or that an MCP connection alone establishes security.

PRIMARY SOURCES

Source ledger: protocols and official engineering documentation

These sources substantiate specific architecture concepts. The article distinguishes documentation-backed facts from SXF analysis.