Skip to main content

Evaluate a plugin

Plugin evaluation uses the same setup, run and results commands as other interfaces. It is currently unreleased; use a development checkout containing plugin support. Evaluation requires Docker, the Harbor/Python runtime and the selected model credential.

Choose the native format

Intermesh preserves the selected bundle and uses the coding agent’s native loader. These are the combinations supported by the evaluator: Every configured agent must support the selected format. Intermesh does not convert bundles between formats. Omit interface.format for detection; choose it explicitly when multiple native manifests are present. A portable root manifest takes precedence. For OpenCode, set interface.entry when a module path is needed instead of the package’s exports/main entry.

Configure the evaluation

Run inter eval setup my-plugin to install the project Skills. Use inter-eval-setup to create .intermesh/evaluations/my-plugin/eval.json, with interface.kind set to plugin. Start with the complete configuration example and replace its illustrative source, agent/model and scenario values. Use a local directory with an explicit include list, a public Git source pinned to a full commit, an exact public npm package, or an npm archive. Paths are relative to eval.json unless absolute. Include runnable files, manifests and required components. Plugin Git sources omit build; preparation installs npm dependencies in Docker without a separate plugin build step. Credentials use environment references in eval.json. Keep values out of selected source files and configuration. Interactive plugin/OAuth sign-in is not automated.

Prepare, run and inspect

Setup records component inventory and checks format compatibility. Dry run checks execution prerequisites. Review the run scope and approve the interactive prompt before trials use model APIs. Prepare again after changing the selected source.

Choose checks that match the task

  • Bundled MCP tools: kind: "mcp"; optional match.server selects a declared server.
  • OpenCode custom tools: kind: "native-tool", with a tool name and input/result expectations.
  • Bundled CLI tools: kind: "cli"; list executable names in interface.commands for recording.
  • Output files: kind: "file", with existence, nonempty or checksum expectations.
  • Call sequences: kind: "call-sequence", with steps referencing existing call requirement IDs.
  • Call budgets: kind: "interface-call-count", combining the plugin’s recorded tool and CLI calls.
Declared skills and hooks are inventory, not proof that they ran. Output checks establish file properties; configured review expectations are agent self-assessment. Neither alone proves component activation. A completed trial can fail its checks, and missing evidence can leave it unassessable. When order matters, add a sequence alongside its call requirements:
Each selected call must succeed and complete before the next starts. Help and retries may intervene. Parallel calls and missing ordering evidence remain unassessable. Mixed CLI/MCP transitions need a shared proof; separate timestamps are not enough. See the sequence grading guide for exact verdict rules and evidence limits.