AI companies publish documents setting out how their models should*
behave. This
grid shows, behaviour by behaviour, how far each one actually goes.
Missing a specification, or a behaviour the index should be asking about?
Propose one.
Loading.
Some AI companies publish a model behaviour specification, a document setting out how their models should* behave. This view scores whether each company publishes one, what it covers, whether changes to it can be seen, and whether the public gets a say before it is weakened. It covers nine companies as of September 2026, and the scores are Polaris Collective's own and open to challenge.
Loading.
Open weights marks a company whose flagship model, or nearly, anyone can download and run. None of these questions can reach how such a model behaves once it has been downloaded, so part of a low score there reflects a limit of the questions.
Meta and xAI tie on 5. Meta is placed sixth on the supporting practices and on its direction of travel: it has committed in writing to publish a model behaviour specification, and it already tests its models against an internal one.
As of September 2026
Having a specification and covering models with it are scored separately. Alibaba's specification names no model, and on a single "has a specification" column it would score the same as one that binds a company's products.
The checks say more than the total. The totals put OpenAI and Anthropic one point apart, and they fall short in opposite ways. OpenAI keeps a record of its versions and does not list its hard constraints, while Anthropic lists its hard constraints and keeps no record of its versions.
Companies whose models anyone can download are marked. Mistral AI, Moonshot AI and DeepSeek take three of the four lowest places, and a reader who sees only the totals might take them for the least responsible companies. These questions have no way to reach a model once it has been downloaded, so the table marks these companies "Open weights".
The bottom of the ranking comes down to the first question. Six of the nine companies publish no model behaviour specification, so most of the grid cannot apply to them and they sit at the bottom for that reason. Among the companies that do publish, the change log is what separates them, while the totals show the case for publishing at all.
A caveat. The four questions were written to argue for change, and turning them into a scored grid gives them more precision than they were designed for. The descriptions of each score are ours, and a different reading could move a company by a few points. The overall picture would stay the same: all nine fall short on the change log, none offers a comment window before a hard constraint is weakened, and none covers government, defence and other special deployments.
The four questions are the four asks of a transparency proposal Polaris Collective has been drafting since 10 September 2026. Each is split into checks scored from 0 to 4. The first two questions have three checks each, because together they are the minimum the proposal asks for; the other two have two. The total is out of 40.
The tables below say what earns 0, 2 and 4 on each check. A score of 1 or 3 falls between the descriptions either side of it.
Five further practices, taken from Kembery and colleagues, Emerging International Best Practices for AI Model Specs, keeping only those an outsider can check from public sources. Each is scored 0, 1 or 2, for a total out of 10. They are reported beside the ranking and left out of its total, because they measure related good practice.
Some trading names mislead, so here are the companies behind three of them. Moonshot AI is Beijing Moonshot AI Technology (北京月之暗面科技有限公司), which makes the Kimi models. DeepSeek is 杭州深度求索人工智能基础技术研究有限公司, of Hangzhou, spun out of the High-Flyer (幻方) quantitative investment fund. Mistral AI is Mistral AI SAS, of Paris.
Researched on 18 September 2026 from primary sources, with one search for evidence per company. Every score rests on a document we found and dated. These points are not yet confirmed.