Automated data collection delivered faster and more reliable trading insights.
Statistics Denmark works with some of Denmark's most sensitive public data. That's why solutions like ChatGPT and Gemini weren't an option from the start.
The project began as an internal research project to explore how AI could be used at scale at Statistics Denmark.
The organization already had strong internal development skills and had invested in its own infrastructure, including newer GPUs for running AI models internally. So the question wasn't whether AI could be used, but how to create solutions that were useful, scalable and secure – without becoming dependent on commercial AI APIs.
Among other things, the project included:
The architecture was built on a simple principle:
Raw data stays internal. Only the necessary vector representations can be used outside.
Internal models on Statistics Denmark's own infrastructure handle anonymization and vectorization of data. External AI models only work with these representations and never get access to the underlying raw data.
This creates a hybrid solution that combines the security of a closed environment with the flexibility and capacity of external AI models.
As part of the solution, BCT also built an internal chat tool based on a large open-source model. It gave employees the ability to work with sensitive data through natural language in a secure environment.
In addition, we set up an MCP server around Statistics Denmark's public API. This lets external AI agents retrieve structured data directly from the API instead of scraping the website.
The result wasn't just new AI tools, but a more scalable foundation for Statistics Denmark's continued work with AI.
We designed an architecture where internal models handle anonymization and vectorization on Statistics Denmark's own infrastructure.
External models only work with the vector representations. That way, raw data stays protected while the organization can use modern AI models where it makes sense.
Statistics Denmark's data tables are large and complex and can be hard to navigate.
We made the data searchable through embeddings and vector databases, so users can search by meaning and context to a greater extent instead of having to know the exact structure or wording in advance.
External AI agents could previously create unnecessary load by scraping the website.
So we built an MCP server on top of the existing API, so AI agents can instead get structured access to public data directly from the backend.
This gives more stable access to data and reduces the load on the user-facing part of the website.
We built a secure internal chat environment based on an open-source model with 120 billion parameters, running on Statistics Denmark's own infrastructure.
Here, employees can work with vectorized, sensitive data through natural language. The solution has also been developed with the option of dynamic visualizations.
The technology was only one part of the project.
Over six months, we therefore ran a training program focusing on AI observability, automation tools, integration into development environments and compliance, among other things.
The goal was for Statistics Denmark to have not just a solution, but also the internal skills to continue working with AI on its own.
Sensitive data could now be vectorized, searched and used in AI-based workflows without leaving the internal infrastructure.
At the same time, external models could be used when more capacity was needed – without getting access to the underlying raw data.
This made it possible to:
More broadly, the project showed how organizations with high security and compliance requirements can work with AI without giving up either control or opportunities.
Automated data collection delivered faster and more reliable trading insights.
We helped Hafslund Kraft move from manual finance processes to a modern, cloud-based solution.
Governance and security for health data at national scale