Agentforce governance and evaluation
Agentforce governance and evaluation
Define what Salesforce agents may know and do, then verify permissions, actions, escalation and release behaviour with repeatable evaluations.
Teams preparing to release or expand an agent that can access business information or perform system actions.
- Topic
- Evaluation set
- Score
- Release gate
- Production
- Owner
Diagram of agentforce evaluation before a production action with labelled stages.
Intended outcomes
What this work should change.
- A reviewed knowledge and action permission matrix
- Repeatable release evidence for normal, adversarial and failure cases
Govern behaviour at the point of use
An agent policy is useful when it changes what the system can actually retrieve, disclose and execute. A document saying that the agent should be careful does not establish a permission boundary. Emerge connects governance decisions to source access, tool contracts, approval requirements and repeatable release tests.
The programme can assess a new Agentforce implementation or strengthen an existing pilot. We inspect the actual configuration, actions and conversation channels rather than relying only on intended behaviour. The goal is a controlled operating scope that the business owner, platform team and reviewers can explain and verify.
Build a knowledge and action matrix
The matrix begins with audience and task. A public visitor, an authenticated customer and an internal adviser may use the same conversational interface while needing different information. We identify the permitted source set and the context required before restricted records can be retrieved.
Actions are classified by their effect. Reading a public article, preparing a case draft and changing a commercial record have different consequences. Each operation records its input contract, executing identity, authorisation check, approval requirement and evidence of completion. The model cannot grant itself a broader role simply because a request sounds plausible.
We also identify forbidden transitions. An unauthenticated conversation must not become an account lookup merely because the user supplies a matching name. A retrieved document must not be allowed to redefine the agent’s instructions. A proposed action must remain distinguishable from a completed destination change.
Turn policies into executable checks
The evaluation set uses representative questions and transactions, with expected behaviour defined before testing. Normal cases establish useful capability. Boundary cases establish what the agent should refuse, clarify or hand over. Failure cases establish whether it reports uncertainty and preserves the user’s next step when a dependency fails.
Adversarial examples include attempts to override instructions, request another user’s information or turn source material into an instruction to invoke a tool. They use safe synthetic content and controlled destinations. The test is whether the deployed controls hold, not whether the agent can recite its policy in response to a direct question.
Tool tests validate parameters independently of conversation quality. Unsupported identifiers, unexpected fields and invalid state changes are rejected at the application boundary. Permission checks run with the identity that will execute the operation, and a denied action remains denied even if the conversational model attempts it repeatedly.
Examine uncertain and partial outcomes
Distributed actions can fail after a destination has accepted work. Blindly retrying such an operation can duplicate a case, order or notification. We require the action contract to identify how an uncertain result is checked and when a human must reconcile it before another attempt.
The agent’s wording is evaluated against the actual system state. It can say that a request was received when receipt is verified, but cannot call a transaction complete while downstream processing remains pending. If the outcome is unknown, the conversation should present that uncertainty and the recovery path accurately.
Approval also needs state discipline. A person approves a defined action with specific parameters and relevant context. A later material change should not inherit approval automatically. The review evidence connects the proposal, approval and executed operation so an operator can understand what happened without reconstructing an entire conversation.
Review traces without creating unnecessary exposure
Investigation requires enough evidence to connect a question, retrieved source, tool attempt and destination result. It does not require unrestricted storage of every sensitive field. We define redaction, access and retention rules for the traces used by reviewers and operations staff.
Reviewers classify failures into useful categories: unsupported answer, wrong source, access error, invalid action, poor handover or technical unavailability. That classification directs the correction to the appropriate owner. A prompt change is not the answer to every failure; some issues require a source update, contract repair or stricter server-side permission.
Metrics distinguish severity as well as frequency. A low-frequency data disclosure cannot be averaged away by many correct public answers. The release gate names unacceptable failures and the evidence required to resolve them before the scope expands.
Release through controlled increments
The rollout begins with a defined topic, audience and action set. Evaluation runs against the exact configuration intended for release, and previous versions remain available for the agreed rollback process. Changes to knowledge, model configuration or tool contracts trigger the relevant regression cases.
A pilot includes a monitored handover to people who can act on unresolved work. The business owner reviews whether the agent completes useful tasks and escalates appropriately. Expansion adds new permissions or actions deliberately, with updated examples and acceptance criteria.
Make governance an operating responsibility
Emerge hands over the permission matrix, evaluation set, release criteria, trace-review procedure and incident runbook. Named owners manage source changes, access decisions and tool maintenance. A recurring review examines new customer questions, recurring failures and changes in the surrounding business process.
The result is an agent whose useful scope can grow through evidence, with controls that live in the implementation and a review process that remains practical after launch.
Your next move
Bring us the operating problem.
We will help you decide whether Agentforce governance and evaluation is the right starting point, what to implement first and who owns the result.