CollectionsField terms42
Vocabulary for building AI agents that plan, act, and recover from failure in production systems.
For developers starting to build with AI agents and coding assistants.
Lobyas brings these back across Read, Review and Quiz until they stick.
Invite-only: request access, or use the code you were given.
Read, review and quiz them in Lobyas, with no due dates to fall behind on.
Invite-only: request access, or use the code you were given.
Tying a model's claims to a verifiable source, like a document or a tool result, instead of letting it answer purely from its own trained knowledge.
Constraints put around an agent's behavior to prevent harmful, out-of-scope, or unintended actions. Can be implemented as input/output filters, tool restrictions, or a separate validator model.
Example:A guardrail blocks any tool call that would delete a database record without a confirmation step, regardless of what the agent decides.
A model confidently generating information that sounds plausible but is fabricated or wrong. In agentic systems this is especially dangerous because the agent may act on false information.
Example:An agent invents a non-existent API endpoint and tries to call it, then fails to handle the 404 and spirals into a retry loop.
A design pattern where the agent pauses at defined checkpoints and waits for a human to review or approve before continuing — rather than running fully autonomously.
Example:An email-drafting agent writes the message and shows it to the user before sending. The user clicks 'approve' or 'revise'.
A property of an action where running it twice with the same input has the same effect as running it once, so a retry after a timeout can't cause double damage.
Example:Charging a customer with an idempotency key means a retried payment request never bills them twice.
Logic that notices when an agent is repeating the same action or getting the same result over and over, so the system can break out instead of burning budget indefinitely.
Example:An agent stuck calling the same failing search query five times in a row gets flagged and stopped rather than trying a sixth.
A pass where an agent reviews its own previous output or action result and decides whether to revise it, retry it, or move on.
Code that automatically re-attempts a failed action a limited number of times, usually with a growing delay between tries, before giving up.
Example:A tool call that times out gets retried after 1 second, then 2, then 4, before the agent finally reports failure.
The rule that tells an agent when to stop looping and return a final answer, whether that's reaching the goal, hitting a step limit, or running out of budget.
Information an agent can store and retrieve across steps or even across sessions. Typically categorized as in-context memory (inside the prompt), external memory (a database), and episodic memory (records of past interactions).
Example:An agent remembers from a previous session that the user prefers metric units, so it doesn't ask again.
Deliberately designing what goes into the model's context — system prompt, retrieved docs, tool outputs, prior turns — rather than treating the prompt as an afterthought.
Example:A coding agent's context engineer trims stale file diffs from the last five turns before every call, instead of letting the transcript grow unchecked.
The repeated cycle an agent runs: observe the current state, think about what to do next, take an action, observe the result, repeat — until the goal is met or a stopping condition is hit.
Example:A research agent loops: search → read page → decide if it has enough info → if not, search again with a refined query.
An architecture where multiple specialized agents collaborate on a task, each handling a different sub-problem, with some orchestrating mechanism routing work between them.
Example:A research pipeline where a Planner agent breaks down a question, a Search agent gathers sources, a Writer agent drafts the report, and a Critic agent reviews it.
A backup model an application switches to automatically when the primary model is unavailable, rate-limited, or returns a bad response.
Example:When the primary model returns a 503, the app silently retries the same request against a smaller fallback model instead of failing outright.
The ability to see what an agent actually did step by step after the fact, through logs, traces, and metrics, rather than just its final answer.
Storing a prefix of a prompt (like a long system prompt or document) on the provider's side so repeated calls skip re-processing it, cutting cost and latency.
An AI system that can autonomously take a sequence of actions to complete a goal, rather than just producing a single response. It decides what to do next based on observations from previous steps.
Example:A coding agent that reads failing tests, writes a fix, runs the tests, and iterates until they pass — without you guiding each step.
A hidden instruction block prepended to every conversation that tells the model its role, constraints, available tools, and behavioral rules. It's the developer's primary way of shaping agent behavior.
Example:A system prompt might say: 'You are a customer support agent. Only answer questions about our product. Never discuss pricing.'
A model output formatted as a structured call to a specific tool with named arguments, instead of plain text describing what it wants to do.
Example:Instead of writing 'let me check the weather', the model returns {"function": "get_weather", "args": {"city": "Almere"}} that your code can execute directly.
Model output constrained to a fixed schema, like a defined JSON shape, so downstream code can parse it without guessing at the format.
The component or step that decides what sequence of actions an agent should take, kept separate from the steps that actually execute those actions.
Breaking a large, ambiguous goal into smaller ordered subtasks that can each be handled with a single reasoning step or tool call.
Example:"Book me a trip to Lisbon" becomes: search flights, compare prices, check the calendar for conflicts, draft the itinerary.
A prompting technique where the model is instructed (or trained) to write out its reasoning step by step before producing a final answer, which tends to improve accuracy on complex tasks.
Example:Instead of directly answering 'What's 17 × 24?', the model writes: '17 × 20 = 340, 17 × 4 = 68, total = 408.'
A prompting pattern where the model alternates between Reasoning (thinking out loud) and Acting (calling a tool). Each cycle of thought → action → observation tightens the agent's approach.
Example:Thought: I need today's weather. Action: call weather_api('Amsterdam'). Observation: 18°C, cloudy. Thought: Now I can answer.
An isolated execution environment where an agent's actions, like running generated code, are contained and can't touch the real system or network.
Limiting exactly which actions and data an agent's tools are allowed to touch, so a bug or bad decision can only cause damage within that narrow boundary.
The maximum amount of text (measured in tokens) an LLM can process in a single call. Everything the agent knows about the current task must fit inside this window.
Example:If an agent's context window is 128k tokens and a codebase has 200k tokens, it can't read the whole codebase at once — it must select relevant files.
A component (often itself an LLM) that manages the overall workflow: deciding which sub-agent or tool to invoke next, passing outputs between steps, and tracking progress toward the goal.
Example:In LangGraph, a supervisor node reads the latest state and routes to either the 'researcher' or 'writer' node depending on what's still needed.
The ability of an LLM to call predefined functions or external APIs as part of generating a response. The model outputs a structured call (e.g. search(query='...')) and the host application actually executes it.
Example:An agent calls a web_search tool, gets back results, then incorporates those results into its next reasoning step.