Explyt / Webinar guide

Working with AI Tools at the User Level

The extended guide to the webinar: methodology, approaches, feedback loops, parallel work and subagents, and personal setup — everything from the session, in one place, ready to apply in your IDE.

  • 5 chapters + key terms
  • 30-minute setup checklist
Sergey Pospelov
Sergey Pospelov · Speaker, Explyt

Questions after the webinar? Message me on LinkedIn — I'm happy to help.

Recap: key terms

The vocabulary used throughout the guide. Everything else builds on one idea: output quality equals context quality.

Harness
LLMthe model Agentloop + tools AssetsSkills · Rules · MCPs
The full setup that does the work. The same LLM gives very different results depending on the agent that runs it and the assets you give it. The model is the part you control least; the assets are the part you control most.
AI agent
An AI assistant with tools. Reads and edits files; runs builds, tests, and terminal commands.
Context
Everything the model sees when generating a response: the system prompt, chat history, files, and tool outputs.
Rules
Customize agent behavior (added to the system prompt).
Skills
Agent know-how for common tasks: prompt + scripts + resources.
MCP servers
External tools for the agent.
AGENTS.md
A project description for AI, stored in the repository.
The golden rule

Output quality = context quality.

Methodology = predictable results

The cost of working without a methodology shows up in three ways.

01 · Direction

The wrong task

The agent confidently solves the wrong task. It invents requirements and relies on assumptions. Hours of work go to waste.

02 · Reliability

Hidden bugs

The code looks correct and passes a quick review. A week later, it breaks in production. Debugging costs more than starting from scratch.

03 · Repeatability

No reproducibility

It works this time, but not the next. You cannot share the practice with your team or scale it.

Response quality depends on context usage

The percentage of context used directly affects the quality of the model's responses.

0%

Completely empty context (never happens)

12%

An empty chat

50–85%

Model response quality declines

99%

The model cannot respond

Context-usage figures reproduced from the webinar presentation.

Three anti-patterns that waste hours

Every anti-pattern has a symptom you can spot and a set of fixes you can apply.

Anti-pattern 1 of 3

Insufficient context

The agent solves the wrong task. It fills in the gaps for you, relies on assumptions, and hallucinates.

Example: you ask to "add tests" — the agent adds unit tests, but you wanted e2e tests. The agent is not to blame: it did not know.

  • Precise instructions. Not "add tests," but "add unit tests for class X."
  • Rules + AGENTS.md. Define project conventions, style, and libraries once.
  • Specification-Driven Development. Write the specification before implementation.
Anti-pattern 2 of 3

Overloaded context

The agent mixes up tasks. It returns to old tasks, mixes them together, and loses focus.

Example: you spend three hours on different tasks in one chat. On the fourth task, the agent applies a pattern from the first to code from the third.

  • One chat — one task. Start a new task? Open a new chat.
  • Large tasks → subtasks. Handle each subtask in a separate chat.
  • Chat compaction. Moving to another task with the same context? Compact the chat if you are sure enough context will remain.
Anti-pattern 3 of 3

Blind trust in the output

The code looks correct. It passes your visual review but breaks in production.

Example: you skipped the tests and trusted what you saw. The agent sounds confident, as though it tested the code. A week later, a bug appears in production.

  • Automated validation. Tests, linters, and scripts — not just your eyes.
  • Human-in-the-loop. Check the agent's work at key points: specifications, plans, tests, and more.
  • Review agent. Always use a separate agent: agents praise their own work. Ideally, use a model from another vendor.

Specifications before implementation. Validation before code.

Idea: solution quality depends on an accurate understanding of the project and a precise task description. So a specification is needed before implementation — a reusable project description and a task description. Writing specifications takes effort, so the agent drafts them in planning mode; you validate and edit.

Approach 01 · Specification-Driven Development (SDD)

Project level

AGENTS.md

A reusable project description, stored in the repository and shared with the team. In Explyt, create and update it with a Skill.

Skill prompt
Create AGENTS.md for the current project
Task level

Planning

Plan mode / Skill.

  1. The agent proposes a plan for the task.
  2. You validate and edit it.
  3. Break it down into files: Plan.md, task-01.md…

Why make specification a separate phase?

During specification, the agent gathers extensive context and summarizes it. During implementation, the agent uses the specification without context overload.

"Writing specs takes too long."

The agent writes them — you just review and refine.

Specification before code

Solution quality = project understanding × task precision.

What a task specification looks like

A sample task-01.md the agent drafts in Plan mode. You review it, fix it, and approve it. Then the agent implements from this file, not from a long chat history.

task-01.md
# Task 01 — Rate-limit GET /api/orders

## Context
- Service: order-service (Spring Boot 3, Java 21), see AGENTS.md
- Source: GitHub issue #482
- Touches: OrderController, ApiKeyFilter

## Goal
Limit each API key to 100 requests per minute on GET /api/orders.

## Requirements
1. Over the limit → 429 Too Many Requests with a Retry-After header.
2. Limits are per API key, not per IP.
3. Configurable: orders.rate-limit.per-minute (default 100).
4. Requests without an API key still return 401 (unchanged).

## Out of scope
- Other endpoints
- Distributed (multi-node) limits

## Design
- New RateLimitFilter, registered right after ApiKeyFilter.
- In-memory token bucket per API key (Bucket4j).

## Acceptance criteria (tests first)
- [ ] 100 requests within a minute → all 200
- [ ] 101st request → 429 + Retry-After
- [ ] Two API keys have independent limits
- [ ] Property override changes the limit
- [ ] ./gradlew check passes (tests + linter)

## Plan
1. Write failing tests: RateLimitFilterTest (MockMvc).
2. Add Bucket4j and RateLimitProperties.
3. Implement RateLimitFilter.
4. Run tests until green, then review (max 3 rounds).
  • ContextLinks the project description, the issue, and the code it touches. The agent doesn't have to guess.
  • RequirementsNumbered and testable, including behaviour that must not change.
  • Out of scopeDraws the line, so the agent doesn't overengineer.
  • Acceptance criteriaThey become the tests you write first. This is where SDD meets TDD.
  • PlanSmall, ordered steps. Each one can go to a separate chat or subagent.

TDD: validation before code

An agent can always make mistakes. It needs a way to self-check before handing work to a human.

  1. 1

    Tests

    The agent writes tests.

  2. 2

    Implementation

    The agent writes code.

  3. 3

    Test run

    The agent runs tests.

Greenor → back to step 2

What should the tests be based on?

The task description in a GitHub Issue

Connect GitHub MCP and attach the issue link.

The specification — the task plan (SDD + TDD)

When planning (Plan mode / Skill), specify that you want to use TDD. Check that the agent schedules test writing before implementation.

The approach: SDD + TDD

SDD controls the input; TDD controls the output. Together they catch errors before a human ever sees them.

Input control · SDD

Input

Creates a high-quality description of the project and task.

Project
AGENTS.md
→
Task
Plan
→
Agent
✳
Output control · TDD

Output

Creates an automated system for validating the output.

Agent
✳
→
Tests
Run
→
Result
Green?

Solving a real task with SDD + TDD

Six steps from issue to merge. You own the checkpoints; the agent does the work in between.

  1. 1You

    Describe the task

    Issue link, goal, constraints. AGENTS.md already gives project context.

  2. 2Agent

    Draft the spec

    Plan mode → task-01.md with acceptance criteria and a plan.

  3. 3You

    Approve the spec

    Fix wrong assumptions now. Mistakes here are the cheapest to fix.

    ✓ checkpoint
  4. 4Agent

    Tests first

    Failing tests from the acceptance criteria. You approve them, then lock them with Edit scope.

    ✓ checkpoint
  5. 5Agent

    Implement until green

    Code → run tests → fix. It repeats until every check passes.

    ↻ loop
  6. 6Agent

    Review, then merge

    A review agent runs at most 3 rounds, then you read the diff and merge.

    ↻ loop
YouAgent✓ Human checkpoint↻ The agent repeats until the check passes
Rule of thumb

Spend your attention on steps 3, 4 and 6. Everything between them is the agent's job.

Methodology spectrum: from vibe coding to SDD + TDD

Task complexity increases from simple / prototype to complex. Escalate the approach only when the simpler one fails.

Task complexityApproachWhat to do
Simple / prototypeVibe codingCreate AGENTS.md, choose the right agent mode and Skill, describe the task in the chat.
Medium / complexSDDThe agent misunderstands the task → add a specification.
Medium / complexTDDThe solution is buggy → add tests before implementation.
ComplexSDD + TDDStill not solved? Combine both.

How much autonomy should you give the agent?

  1. 1

    Manual control

    Approve every step.

  2. 2

    Plan + Execute + Review

    Validate the plan, the review, and the code itself.

  3. 3

    Full autonomy

    Validate only the output.

The key principle

Greater autonomy requires better context and stronger validation.

To get good results, give the agent a clear definition of "good"

The agent runs the checks itself and keeps working until the criterion is met.

Define the success criterion
Keep writing code until all tests in the e2e folder pass

Example criteria

Anything the agent can check on its own and report back.

  • The linter passes
  • Tests pass — unit, integration, and automated tests
  • No compilation errors or warnings (Explyt checks automatically)
  • Test coverage exceeds the threshold (Explyt: Increase coverage)
  • The UI works in simple scenarios (Playwright MCP, Chrome MCP)

Anti-pattern: tailoring tests to the implementation

Tests are green but check nothing. The agent edits tests instead of code: weakens assertions, uses mocks, or marks tests as skipped. Example: a test was failing, so the agent disabled it — the criterion is technically met, but the bug remains.

  • Prevent test editing: block the test directory with .agentignore (Explyt: quick "Edit scope" settings)
  • Separate roles: one agent / chat writes tests; another implements the code
  • Review test diffs: a separate Review agent checks test changes first, then you do

A review loop with a hard stop: at most 3 rounds

Tests catch broken behaviour. A review agent catches the rest: missed requirements, risky code, weak tests. Let the agent fix what the review finds, but cap the number of rounds.

Prompt
Implement the task from task-01.md. When all tests pass, run a review subagent on your diff. If it reports problems, fix them and run the review again. Do at most 3 review rounds. Stop when the review is clean or after round 3, then report: what you fixed, what is still open, and why.
Step 1ImplementUntil the tests pass
Step 2ReviewSeparate subagent, fresh context
Step 3FixOnly the reported problems
Round 1 · 2 · 3 max
Step 4ReportFixed · still open · why

Why a separate reviewer?

Agents praise their own work. A fresh context, or a model from another vendor, sees what the author missed.

Why stop after 3 rounds?

Without a cap, the agent can loop forever on nitpicks, burn tokens, or start rewriting good code. Anything still open after 3 rounds needs a human decision.

Specialization, clean context, and independent work in parallel

A subagent is a separate agent with its own context, system prompt, and tools.

01

Specialization

By task type: review, testing, and code search.

02

Clean context

The subagent returns only the result, without cluttering the main chat.

03

Parallel work

Independent tasks run in parallel.

Do not overcomplicate things
"We have seen teams spend months building complex multi-agent architectures, only to discover that better prompting of a single agent delivers the same result."
Anthropic · as quoted in the presentation

Subagents: when to use them

When to use

Large, independent work

Work that splits cleanly into separate steps.

Large features
Many files and a complex pipeline of separate steps.
Project-wide analysis
Gathering documentation and finding functionality across the project.
Review pipelines
Checking code against rule sets, with independent checks in parallel.
When not to use

Small tasks or close supervision

A single agent is the better default.

Small and medium tasks
One agent is faster and cheaper.
The task fits in one chat
Bug fixes, small features, and improvements to code written in the chat.
You need control over every step
Multiple subagents are harder to supervise.
Higher token usage
Each subagent gathers context from scratch.

Advanced AI agents decide when to launch subagents (Claude Code, Cursor, Explyt in general mode). But the system is not perfect: it may launch too many or too few. Set boundaries and preferences with Rules and Skills.

Parallel work: Git worktrees

git worktree — multiple working copies of one repository, each on a different branch. Scenario: each working copy has its own agent. One agent builds feature A, another builds feature B, while you review a bug fix.

⌥ One repository
Working copy AAgentBuilds feature A
Working copy BAgentBuilds feature B
Working copy CYouReview a bug fix
Limitations

Realistically, you can supervise 2–3 agents at once. Your review is the bottleneck, not code generation.

Personal preferences. Shared project conventions. Project memory.

Agent settings live in two places: your home directory, and the repository.

Personal · yours

Home directory

Not added to the repository.

  • Global Rules — response language, communication style, and degree of autonomy.
  • Personal Skills — tasks and approaches not tied to a project.
Project · in the repository

Shared with the team

Version-controlled and committed.

Three kinds of assets, three jobs

Assets are the part of the harness you control. Each one enters the agent's context differently. Pick the kind that matches the problem.

How to behave

Rules

Standing instructions

Always in context

Added to the system prompt, for files that match their pattern.

“Answer in English. Use JUnit 5 and AssertJ. Never edit generated code.”

  • Global: personal style for every project
  • Project: team conventions, committed
How to do a task

Skills

Know-how on demand

Loaded when relevant

The agent sees only the name + description and opens the full SKILL.md when a task needs it.

“Review a DB migration against the team checklist.”

  • Global: ~/.explyt/skills
  • Project: .explyt/skills/<name>/SKILL.md
What it can reach

MCPs

External tools

Called as tools

The agent calls them to read or act outside the IDE; only the tools you allow.

“Read issue #482 from GitHub” · “Open the page in Playwright.”

  • Global: ~/.explyt/mcp_servers.json
  • Project: .explyt/mcp_servers.json
Which one?

The agent should always follow it → Rule. It's a procedure for some tasks → Skill. It needs data or actions outside the IDE → MCP. Personal → global scope; team → project scope.

Project memory

Memory Bank is project memory: durable, non-obvious knowledge about this project — conventions, pitfalls, decisions — stored as Markdown in the project's .explyt/memory/ folder. There is no personal memory that follows you from project to project: keep your own preferences in global Rules and personal Skills.

  1. 1

    Markdown files

    An index (MEMORY.md) plus one file per entry in .explyt/memory/. Readable and editable by hand.

  2. 2

    Recall

    The agent sees the index and opens the entries relevant to your request.

  3. 3

    Update

    It adds new entries and merges or removes outdated ones.

Goal: capture project knowledge once — where the agent can find it next time — instead of re-explaining it. Keep personal preferences in global Rules. In Explyt, Memory Bank was introduced in release 5.14.

30-minute setup

0 / 4 done
Checklist progress is saved in this browser only.
Explyt shortcut · 5 min

Just run Onboard the Agent to Your Project. The agent asks about your preferences and proposes ready-to-use Rules, Skills, and AGENTS.md.

Onboarding prompt
Onboard the Agent to Your Project

Let's keep the conversation going

If you have questions about the webinar, want to share how the approach worked for your team, or got stuck setting up your agent — write to me on LinkedIn. I read every message.

Sergey Pospelov
Sergey Pospelov
Speaker · Explyt Team

Sergey covers the practical side of working with AI coding agents in Java and Kotlin projects — methodology, specification-driven workflows, and the feedback loops that keep generated code trustworthy in IntelliJ IDEA and the full JetBrains IDE family.

Ask me a question on LinkedIn

Thank you. Now try it in your IDE.

Explyt runs inside IntelliJ IDEA and the full JetBrains IDE family. IDE-native debugging, semantic refactoring, test coverage feedback, and agent onboarding — built for Java and Kotlin teams.

Try Explyt in your IDE

IDE-native debugging, semantic refactoring, and AI agent onboarding — built for Java and Kotlin teams in JetBrains IDEs.

Download