Building Portable, Governed Skill Packages for Microsoft 365 Copilot Agents

How to package agent skills as self-contained folders, upload them to Agent Builder, govern them at the tool calls that load them, and check them in CI.

TL;DR

  • A skill is a folder with a SKILL.md (name, description, instructions) plus optional resources and scripts. Agent Builder takes up to eight per agent, each a.zip of up to 50 MB.
  • Until a skill is needed, its name and description are all the model sees, so write them as routing rules.
  • Govern skills at the tool calls that load and run them, not in the prompt, and check every package in CI.

Why monolithic prompts break at scale

Put every procedure your agent knows into one system prompt and you pay for all of it on every turn. Google's agent governance docs name the failure: exposing an agent to a very large number of tool descriptions at once causes context bloat, which raises token cost and latency and can degrade the model's reasoning.

There is a second cost. Your business rules now live inside text tuned for one model, and the model changes under you: Microsoft 365 Copilot declarative agents moved to GPT-5.1 with automatic model selection. A rule buried in a prompt has to be re-tested after every such change, and it has no version history of its own.

Skills fix both by treating instructions as lazily loaded modules. Microsoft's Agent Framework describes four stages: advertise each skill's name and description (about 100 tokens per skill), load the full SKILL.md when a task matches, read resource files only when needed, and run bundled scripts only when asked. It also keeps three contracts apart: the skill document format, the MCP transport that retrieves it, and the host SDK's tool registration. That separation is what makes a package portable across models and hosts.

What goes in a portable skill package

A skill is a directory with a required SKILL.md plus optional resource files and scripts. Agent Builder accepts up to eight skills per agent, each a compressed.zip of up to 50 MB, with instructions under 20,000 characters. Agent Framework's guidance is tighter: keep SKILL.md under 500 lines and move reference material into separate files.

expense-report/
├── SKILL.md
├── scripts/
│   └── validate.py
└── references/
    └── policy.md
---
name: expense-report
description: Check an expense report against travel policy before submission. Use for receipts and per-diem limits. Do not use for payroll or reimbursement status.
---
### Instructions
1. Confirm every line item has a receipt attached.
2. Compare totals with the limits in references/policy.md.
3. Run scripts/validate.py with the line items as JSON and report its result.

Spend your effort on the description. The agent picks a skill from its name and description alone, so treat the description as routing metadata, not documentation: say when to use the skill and when not to. If a skill fires too often, the description is too broad; if it never fires, it is too narrow.

Uploading to Agent Builder, and where it will not work

Custom skills in declarative agents are in preview, for organisations in the Microsoft Frontier Preview. In Agent Builder, open Configure, expand Skills, select Add and upload the complete.zip. Microsoft says not to upload SKILL.md on its own. Then test it in Preview with a prompt that should trigger it.

Know two limits first. Custom skills are not available in tenants that use Microsoft Purview Information Barriers, for admin-deployed and user-created agents alike. And Agent Builder hosts the runtime for you, so you cannot put your own policy engine in front of it. Microsoft's route to stronger governance inside Microsoft 365 is copying the agent into Copilot Studio; for full runtime control, host the skills yourself with Agent Framework.

Governing skills outside the model

In Agent Framework, every stage after advertising is a tool call: load_skill, read_skill_resource, run_skill_script. Because loading and running skills are tool calls, a policy layer can govern the whole lifecycle without trusting anything the model writes. Agent Framework leans this way already: all three skill tools require approval by default, and its SubprocessScriptRunner is documented as demonstration-only, with sandboxing, CPU, memory and timeout limits, script allow-lists and audit logs recommended for production.

For policy as code, Microsoft's open-source Agent Governance Toolkit installs with pip:

pip install 'agent-governance-toolkit[full]'

Its Agent OS engine intercepts every agent action before execution at under 0.1 ms p99 and accepts YAML rules, OPA Rego or Cedar. Treat model output as untrusted input: never hand it to a shell or a file path, and validate it against a schema first.

Next steps: gate every package in CI

This script fails the build when a package breaks a documented limit:

import sys
import zipfile
from pathlib import Path
 
MAX_ZIP_BYTES = 50 * 1024 * 1024
MAX_INSTRUCTION_CHARS = 20_000
MAX_LINES = 500
 
def check(package):
    errors = []
    if package.stat().st_size > MAX_ZIP_BYTES:
        errors.append("package is over 50 MB")
    with zipfile.ZipFile(package) as z:
        if "SKILL.md" not in z.namelist():
            return errors + ["SKILL.md must sit at the root of the zip"]
        text = z.read("SKILL.md").decode("utf-8")
    lines = text.splitlines()
    if not lines or lines[0].strip()!= "---" or "---" not in lines[1:]:
        return errors + ["SKILL.md has no YAML front matter"]
    end = lines.index("---", 1)
    meta = dict(l.split(":", 1) for l in lines[1:end] if ":" in l)
    for key in ("name", "description"):
        if not meta.get(key, "").strip():
            errors.append(f"front matter has no {key}")
    body = "\n".join(lines[end + 1:])
    if len(body) >= MAX_INSTRUCTION_CHARS:
        errors.append(f"instructions are {len(body)} characters")
    if len(lines) >= MAX_LINES:
        errors.append("SKILL.md is 500 lines or more")
    return errors
 
if __name__ == "__main__":
    problems = check(Path(sys.argv[1]))
    for p in problems:
        print("FAIL", p)
    print("OK" if not problems else f"{len(problems)} problem(s)")
    sys.exit(1 if problems else 0)
cd expense-report && zip -qr../expense-report.zip. && cd..
python check_skill.py expense-report.zip   # prints OK, exits 0

Then wire up the rest. Log each request, the agent's plan and every tool call, using MLflow Tracing or your existing tracer. Where you host the agent, pin model versions and run regression tests often, because provider updates shift behaviour. Watch which skills fire, and tighten descriptions that misfire.

Sources

copilotagentskillsgovernancebackend

All writing