A new index from Stanford that measures how much foundation models reveal about themselves puts Meta’s Llama 2 at the top of the list — but the big-picture finding is plain: most models remain opaque and none come close to full transparency.
What the index checks
The Foundation Model Transparency Index looks at multiple practical dimensions of openness rather than a single headline metric. The items it checks include:
- Training data: how much information a developer publishes about the datasets used to train a model. This matters because the makeup of training data affects bias, copyright risk and what the model knows. - Compute: whether the creator discloses the compute resources and training steps used. Compute details help researchers replicate results, estimate environmental impact, and judge the scale of a model’s capabilities. - Labour: disclosure about human work involved, for example whether people labelled or filtered training material, and what protections those workers had. Labour transparency touches on ethics and consent. - Downstream use: information about what the model is allowed to be used for, and what safeguards or restrictions are in place for third-party deployments.
The index assesses how much companies and research groups publish in each of those areas. According to the report, Llama 2 comes out ahead of other high-profile models, while several widely used models — including versions of GPT — score lower on the transparency checklist.
Why transparency matters for users and developers
Transparency isn’t just an academic nicety. For users, it affects trust and safety. When you know what went into a model and how it’s intended to be used, it’s easier to judge reliability, spot potential biases, and make informed choices about which services to use for sensitive tasks.
For developers, clear documentation about data and compute makes it possible to reproduce results, compare models fairly and build safer systems on top of them. For policymakers and auditors, transparency is a prerequisite for meaningful oversight: you can’t audit what you can’t see.
Labour transparency also matters because large models are created and cleaned by people. If companies report who did that work and under what conditions, it helps surface ethical concerns and pushes for better practices.
A final practical point: transparency is necessary but not sufficient. An open account of training data doesn’t guarantee a model is fair, accurate or secure. Likewise, a model that is more open can still produce harmful outputs. The index evaluates how much information is available, not the model’s behaviour.
Bottom line: it’s encouraging to see some models being more open, and Llama 2 leads that pack according to the new index, but the overall message is that the industry still has a long way to go before most foundation models are transparent enough for confident public use. Users, builders and regulators will need to keep pushing for clearer documentation and responsible disclosure as these systems grow more powerful.
This article has been restored to the What's New On The Net archive as part of the site's relaunch.