Open source collaboration platforms are adding assistant capabilities, and the plugin frameworks themselves are the least interesting part of the news. What matters is what happened in the four weeks before them. In mid-April an openly licensed instruction-following model arrived with terms permitting commercial use. Last Friday another was released on similar terms, trained from scratch and explicitly positioned for commercial deployment. Others have appeared alongside them. Until this spring, an organisation that could not send its messages to a third-party endpoint had no realistic assistant option. The capable models were hosted, and the weights that could be downloaded were encumbered by research-only licences that legal departments correctly refused to sign. That constraint has loosened, quickly, and self-hosted collaboration platforms are the first place it becomes visible.
The question stopped being whether you can run a model yourself. It became whether the model you can run is good enough for the job you have
That is a considerably more useful question, and it has different answers for different tasks. Treating it as a single yes or no is how organisations end up buying accelerators to run something disappointing.
Match the task to the tier
Mechanical text work. Summarising a long thread, extracting decisions and action items, rewriting a message for clarity, drafting a status note from bullet points. The openly licensed models available today handle this adequately. The quality difference against a frontier hosted model exists but rarely changes the outcome, because a person reads the result anyway. Grounded answers over your own content. Answering a question from your runbooks, policies or channel history. Here the binding constraint is retrieval, not generation. A modest model with well-built retrieval beats a much larger model given the wrong three documents, and most disappointing pilots in this category failed at retrieval while blaming the model. Open-ended reasoning and drafting from nothing. Producing a coherent analysis, reasoning through an unfamiliar problem, writing something genuinely good from a one-line prompt. The gap is real and currently large. Pretending otherwise is the fastest way to lose the internal argument, because the first person to try it will notice. Pilot the first two. Be honest that the third is not yet where the hosted frontier is.
Read the licence before you build on it
Open is carrying a great deal of weight in this discussion. Several widely discussed sets of weights are released under research-only terms, or with acceptable-use restrictions, or under bespoke licences that are not open source in any conventional sense. A prototype built on the wrong weights is a prototype you cannot ship. Check the licence at selection time, record which model and which version produced what, and keep that record. The provenance question will be asked later, most likely by a customer's procurement team rather than by your own lawyers.
Build the abstraction, not the commitment
The release cadence this spring should be treated as information about the future: whatever you select in May will be superseded before the end of the year. Design accordingly. Put a gateway in front of the model that exposes one internal interface, so the model behind it can be swapped without touching the collaboration platform. Keep retrieval as a separate component with its own evaluation, because it will outlive several models. Log prompts and outputs from the first day, both for evaluation and because you will want to compare candidates on your own traffic rather than on a benchmark. The abstraction costs a few days of engineering and converts an irreversible platform decision into a reversible one.
Start narrow, because the chat log is the corpus
A collaboration platform holds the most sensitive continuous text corpus in most organisations: salary discussions, incident response, legal questions, commercial negotiation, the channel where two executives disagree. The temptation on day one is to index everything and demonstrate a company-wide assistant. Do the opposite. Choose one category of channel, make participation explicit, restrict the initial capability to summarisation of content the requester can already read, set a retention period for prompts and generated output, and publish what is being processed and by what. A narrow deployment that people trust expands. A broad one that surprises somebody gets switched off by the general counsel, and does not come back for a year.
Practical Guidance for Mattermost AI Integration
- Classify tasks into the three tiers before selecting any model.
- Fix retrieval first; most model complaints are retrieval failures.
- Verify the licence permits your actual commercial use.
- Put a gateway between platform and model so the model stays swappable.
- Log prompts and outputs from day one for evaluation.
- Start with one channel category and explicit participation.
- Evaluate on your own Arabic and English content, not on benchmarks.
- Set retention for prompts and generated text, not just for messages.
The Regional Angle
Three factors make this development more consequential here than in most markets. The first is that the region has an unusually high concentration of organisations for which hosted assistants were never an option. Government entities, defence-adjacent contractors, national energy companies and the operators of critical infrastructure work under in-country processing conditions, and some run genuinely disconnected environments. For those organisations the arrival of commercially licensed open weights is not a cost optimisation, it is the first time an assistant has been permissible at all. The procurement conversation changes shape accordingly: instead of negotiating processing terms, retention commitments and subprocessor lists with a foreign vendor, the questions become accelerator supply, a supported distribution, and who holds the operational responsibility at three in the morning. That is a different purchase, evaluated by different people, and the organisations that will move fastest are the ones that already run their own infrastructure competently. The second is linguistic and quantitative rather than cultural. The openly licensed models released this spring are trained overwhelmingly on English text, and their Arabic performance is materially weaker — not subtly, but to the point of changing which tier a task belongs in. There is also a cost dimension that rarely appears in vendor material: Arabic script consumes substantially more tokens than equivalent English for most common tokenisers, which shortens the usable context window and raises the cost per summarised conversation. A thread that fits comfortably in English may not fit in Arabic. Before committing to hardware or to a model, run an evaluation on a representative sample of your own bilingual channels, count tokens as well as quality, and accept that Arabic-heavy workloads may require a larger model, a different one, or a hosted fallback for that subset. The third is operational economics that regional buyers consistently get wrong in both directions. Accelerator supply remains constrained globally, lead times are long, and importing, powering and cooling this hardware is a real project rather than a purchase order. More importantly, the economics invert with utilisation: a hosted endpoint costs money only when used, while an owned accelerator costs the same whether it processes ten thousand requests a day or forty. A deployment sized for a pilot and used by twenty people is an expensive way to summarise threads. The staffing point follows from the same logic. A self-hosted inference stack maintained by one capable engineer is a continuity exposure in a market where technical staff move between countries on short notice, so prefer a supported distribution or in-country managed hosting over an arrangement that depends on one person remaining employed.
The objection worth taking seriously
The strongest objection is that self-hosting is a distraction dressed as sovereignty. The quality gap between openly licensed weights and the hosted frontier is wide and, on current trajectories, widening. Hosted vendors now offer enterprise terms that exclude training on customer data, with contractual retention limits and regional infrastructure. Running your own inference means buying scarce hardware, hiring scarce skills and accepting a measurably worse assistant, in exchange for a control benefit that a well-negotiated contract already provides. For most commercial organisations that is a poor trade, and the engineering effort would produce more value applied to almost anything else. This is correct for the majority of businesses, and the honest advice to a typical mid-market company this month is to buy rather than build. It is not correct for the minority that cannot transmit the data at all, whose number in this region is larger than the global average, and for whom no contractual term substitutes for the data never leaving. It also understates the value of optionality. The gateway and retrieval layer described above are worth building even for an organisation that intends to use a hosted model, because they are what allow a switch when pricing changes, when terms change, or when a regulator changes its mind. Build the abstraction; defer the hardware until a task genuinely requires it.
Common Questions
Can a self-hosted model replace a hosted one for everyday work?
For summarising, extracting and rewriting, largely yes today. For open-ended drafting and reasoning, not yet.
Do we need to fine-tune on our own data?
Usually not first. Retrieval over your documents solves most of what fine-tuning is expected to solve, costs far less, and is easier to correct when the underlying content changes.
What hardware does a pilot need?
Less than vendors suggest for a small user population, and the sizing question should follow a measured evaluation rather than precede it. Rent capacity for the pilot before buying any.
What should we expect over the next twelve months?
Expect a rapid cadence of openly licensed releases to continue, narrowing but not closing the gap to hosted frontier models. Expect hosted vendors to respond with stronger enterprise terms and clearer processing commitments, which will reduce the number of organisations that genuinely need to self-host. Expect more employers to restrict staff use of consumer chatbots, as a major electronics manufacturer did last week after confidential material was pasted into one, and expect those restrictions to create demand for a sanctioned internal alternative. And expect the decisive capability difference between deployments to be the quality of retrieval rather than the choice of model.
Mattermost AI Integration — we evaluate open and hosted models on your own bilingual content, build the gateway that keeps the choice reversible, and start with a scope your general counsel will sign.
