Chat-based assistants arrived in collaboration platforms as summarisers. They read the channel you missed, condensed the thread, answered a question about a document. Useful, unthreatening, and fundamentally passive. The self-hosted platforms have now moved past that: agents in Mattermost run inside workflows, act on the operator's own infrastructure and against the operator's own models, and execute steps rather than describing them. That is a different proposition, and the interesting part is not the capability. It is that the whole thing runs somewhere you control, which changes who can adopt it.
An assistant that summarises your channel is a convenience. An agent that executes a workflow step in your channel is an operator, and it needs to be governed like one
Here is what the capability actually is, and what to do before you switch it on.
What is genuinely on offer
Multiple distinct assistants rather than one, each configurable to a purpose and a permission scope, which matters because a single general assistant with everyone's access is the wrong unit of control. Semantic search that respects existing permissions, so retrieval returns what the asking user is entitled to see rather than what the index contains. This sounds like a detail and is the single most important property in the set. Summarisation of threads, channels and meetings, plus image analysis, all standard by now. And the underlying model as a deployment choice — any local or cloud model, including one running entirely on your own hardware. That is the part self-hosted platforms can offer and hosted ones structurally cannot.
Why running agents in workflows is the real change
A summariser produces text a human reads. A workflow agent produces an action a system receives. Once an agent is a step in a workflow, three questions arrive that never applied to the assistant version: what identity it acts under, what it is permitted to do when the triggering message is wrong, and how the outcome is recorded. Most teams enabling this will not ask those questions, because the feature arrived in a platform update rather than in a project.
The practical risk
The triggering content is untrusted. A message in a channel, a document a supplier attached, a ticket a customer wrote. An agent that reads that content and then takes an action is a path from outside text to internal effect, and the mitigation is architectural rather than clever — the content should not be able to authorise the action.
Bound each purpose
Choose the model location, agent identity and narrow permission scope.
Test retrieval
Check restricted content with an unauthorised user and test both working languages.
Keep approval separate
Do not let externally supplied content authorise consequential actions.
Observe and recheck
Retain an operations-visible action log and review behaviour after updates.
Qualitative summary of this article's source text, not a measured outcome or performance estimate.
Practical Guidance for Deploying Mattermost Agents
- Create separate agents per purpose, not one general one.
- Verify that retrieval respects permissions with a real test.
- Decide which model and where it runs before enabling.
- Give workflow agents their own identity, not a person's.
- Keep externally triggered actions on a separate approval path.
- Log agent actions where operations can see them.
- Start with retrieval and summarisation; add execution deliberately.
- Review after platform updates; capability changes without a project.
The Regional Angle
The first regional reason this matters is that self-hosted deployment resolves the data location question outright, and for a significant set of Gulf organisations that is the difference between adopting this capability and not. Government entities, defence-adjacent contractors, and financial institutions under residency expectations cannot route internal conversations through a foreign hosted service, and until recently their practical choice was to forgo assistant capability entirely. Running the platform and the model inside the country removes the objection rather than mitigating it. The second concerns the model choice, which in this region is a procurement and policy decision as much as a technical one. National artificial intelligence programmes in Saudi Arabia and the Emirates have produced locally developed and locally hosted models, and the ability to point a collaboration agent at a model of your choosing means the sovereignty position and the capability decision can be taken separately. Organisations should establish which models are acceptable to their regulator before selecting the workflow use cases. The third is about Arabic, where self-hosting cuts both ways. Choosing your own model means you can select one with genuinely competent Arabic handling rather than accepting whatever a hosted vendor ships, which is a real advantage. But it also means the quality is now your responsibility, and summarisation or retrieval that performs poorly in Arabic will be attributed to your deployment rather than to a vendor. Test retrieval and summary quality in both working languages before rollout, with real channel content rather than samples.
The objection worth taking seriously
The strongest objection is that self-hosting the model is a false economy for most organisations. Running competitive models on your own infrastructure requires hardware, operational expertise and a continuous upgrade commitment, and the capability you get will lag what a hosted frontier model delivers by a noticeable margin. Teams end up with a private assistant that is slower, weaker and more expensive per query than the commercial option, adopted primarily so that a compliance box can be ticked — and users quietly go back to the public tool on their phones. The capability gap is real and the shadow usage pattern is well documented. What the objection misses is the actual decision most regional organisations face. It is not self-hosted versus hosted frontier; it is self-hosted versus nothing, because for the entities with residency constraints the hosted option was never available. Against that baseline, a competent local model doing permission-aware retrieval across two years of internal channels is a large gain regardless of how it benchmarks. The architecture also keeps the choice open — the same deployment can point at a cloud model for non-sensitive workloads and a local one for the rest, which is a distinction no single-vendor hosted assistant can express. Organisations without the constraint should indeed weigh the capability gap seriously.
Common Questions
Should agents execute workflow steps or only draft them?
Start with drafting for anything with an external effect. Execution is appropriate where the action is reversible and the trigger is internal.
Does permission-aware retrieval actually work?
Test it. Create a restricted channel, ask an unauthorised user a question that only that channel answers, and verify. Do not accept the claim on documentation.
What model should we run?
The one that satisfies your residency position and performs acceptably in your working languages, in that order. Benchmarks are a poor proxy for either.
What should we expect over the next twelve months?
Expect self-hosted platforms to keep closing the capability gap on workflow execution. Expect locally hosted regional models to appear in more procurement shortlists. Expect the first governance surprise to involve an agent acting on externally supplied content. And expect permission-aware retrieval to become a standard evaluation criterion rather than a differentiator.
Deploy Mattermost Agents — we set up the agents, the model and the permission boundary together, so execution arrives governed rather than by update.
