Bipko Digital News & Media Platform

collapse
Home / Daily News Analysis / Portugal open-sources Amália, its first national AI model, in a bet on European Portuguese

Portugal open-sources Amália, its first national AI model, in a bet on European Portuguese

Jul 01, 2026  Twila Rosenbaum  42 views
Portugal open-sources Amália, its first national AI model, in a bet on European Portuguese

Portugal has taken a major step toward digital independence with the release of Amália, the country's first national artificial intelligence model built specifically for European Portuguese. The project, officially named Automatic Multimodal Language Assistant with Artificial Intelligence (Amália), is fully open-source, including its weights, training datasets, and source code. This deliberate openness signals a strategic bet on sovereignty, transparency, and cultural specificity in the rapidly evolving landscape of large language models.

Amália is not intended to compete with consumer chatbots like ChatGPT or Google Gemini. Instead, it is designed as a foundational layer for public-sector applications, government services, and academic research. The model is built on EuroLLM-9B, a European foundation model developed through cross-border collaboration. A team of over 60 researchers and students from five Portuguese universities expanded that base with European Portuguese linguistic data, a larger context window, enhanced safety and evaluation mechanisms, and the ability to process images alongside text.

Why open-source matters for a national AI model

The decision to release Amália under an open license is both ideological and practical. On one hand, it aligns with the European push for transparent, auditable AI systems. On the other, it addresses a critical need for a government that intends to integrate AI into citizen services, naval decision-support tools, and other sensitive domains. An open model allows independent verification of training data, bias auditing, and customization – all of which are difficult or impossible with proprietary systems like GPT-4 or Gemini.

Portugal’s Foundation for Science and Technology (FCT) coordinated the project, which also involves NOVA University Lisbon, Instituto Superior Técnico, and the universities of Porto, Minho, and Coimbra. The initial funding of €5.5 million comes from Portugal’s Recovery and Resilience Plan, a European Union-backed initiative. Financing has already been secured through the end of 2027, indicating a long-term commitment rather than a one-off experiment.

The cultural and linguistic edge

The model’s name is a deliberate homage to Amália Rodrigues, the iconic fado singer whose voice is deeply intertwined with Portuguese identity. This cultural reference underscores the project’s ambition: to capture the nuances of European Portuguese, which differs significantly from Brazilian Portuguese in grammar, vocabulary, idiom, and cultural references. Major commercial models are overwhelmingly trained on Brazilian Portuguese data, often flattening these differences. For a public service that needs to speak to citizens in their own register – not an approximation – Amália offers a distinct advantage.

The linguistic gap is more than academic. European Portuguese uses different second-person pronouns, verb conjugations, and prepositions. Words like "autocarro" (bus) versus "ônibus," or "comboio" (train) versus "trem," can create confusion. More subtly, idioms and humor vary. A model trained on Brazilian data might misinterpret a Portuguese citizen’s query about "pão de ló" (a traditional sponge cake) or "pastel de nata" (custard tart) as unrelated to Brazilian equivalents. Amália’s training datasets specifically include European Portuguese sources, from parliamentary transcripts to news articles and literary works.

Technical specifications and future applications

Amália is a multimodal model, meaning it can process both text and images. This opens up use cases such as virtual museum guides that can analyze a photograph of a painting and provide historical context, or AI teaching assistants that can evaluate handwritten student work. The model also includes robust safety guardrails, particularly important for military or citizen-service deployments. Evaluation systems were built from the ground up to measure performance on Portuguese-specific tasks, such as legal document summarization or administrative form assistance.

Planned applications are diverse. The Portuguese Navy is exploring decision-support tools that could analyze satellite imagery and reports in European Portuguese. Museums and monuments are developing virtual guides that can adapt to visitor questions in real time. Citizen services aim to deploy digital assistants that help with paperwork, tax filings, or health queries. In education, an AI teaching assistant could provide personalized tutoring, especially in remote or underserved areas.

A test version of Amália was completed in September 2025 and presented at the PROPOR conference in Brazil, the premier event for Portuguese-language natural language processing. Feedback from the research community helped refine the model’s performance on tasks like machine translation, sentiment analysis, and question answering.

European context and sovereignty concerns

Amália is the latest in a series of European efforts to reduce reliance on American and Chinese AI infrastructure. It follows the OpenEuroLLM alliance, a cross-border initiative to train open models on the continent’s own languages. That project, in turn, builds on earlier work by groups like Aleph Alpha and Mistral AI. Portugal has also seen infrastructure investments, such as Nscale’s €695 million data centre project in collaboration with Microsoft. However, critics argue that renting GPUs by the hour creates the illusion of sovereignty rather than the substance of it. Amália’s open-source nature, combined with its localization focus, aims to address that criticism by ensuring that the model itself remains under domestic control.

The broader European anxiety is plain: language models trained primarily on English, Chinese, or even Brazilian Portuguese data will inevitably reflect the cultural assumptions of those data sources. For a country of just over 10 million people, relying on foreign models for public services carries risks of bias, miscommunication, and loss of linguistic heritage. Amália is a bet that a smaller, more precise model can outperform larger, generic ones in the contexts that matter most for citizens.

Challenges to adoption and long-term viability

The hardest test for Amália will not be technical performance but adoption. Publishing a model openly is one thing; getting universities, startups, and government departments to actually build on it is another. Most sovereign AI projects – from France’s Le Chat to Spain’s ALIA – have struggled to gain traction beyond academic circles. Often, the issue is lack of documentation, limited compute resources for fine-tuning, or the sheer convenience of existing commercial APIs.

Portugal has attempted to mitigate these risks by embedding Amália into existing institutional structures. The five universities involved will serve as hubs for training and support. The FCT is funding fellowship programs for developers and researchers. The government has earmarked pilot projects in three ministries. Whether that is enough remains to be seen, but the 2027 funding horizon provides at least a runway for experimentation.

Another challenge is compute. Training and running large language models requires substantial GPU capacity, which is still scarce and expensive in Europe. While Amália’s 9-billion-parameter size is modest compared to GPT-4’s estimated trillion-plus parameters, it still demands significant resources. The project so far has relied on cloud credits and national supercomputing infrastructure. Long-term sustainability may require either a dedicated EU cloud for open-source AI or cost reductions in hardware.

Privacy and data governance also loom large. Because Amália is fully open, any organization can inspect its training data – and potentially identify sensitive information if de-anonymization techniques are applied. The project has implemented data filtering and redaction pipelines, but the risk is inherent in open models. For military applications, the Navy may need to fine-tune its own closed version, which is legally possible under the open license but would require additional resources.

Despite these hurdles, Amália represents a concrete attempt to build AI infrastructure that reflects a specific linguistic and cultural identity. It avoids the trap of trying to compete with large commercial models on general benchmarks, instead focusing on niche excellence. If successful, it could serve as a template for other small language communities – from Catalan to Finnish to Welsh – that face similar pressures of homogenization by dominant AI systems.

The model is now available for download from the project’s repository, along with documentation and evaluation scripts. The open licence allows commercial and non-commercial use, modification, and distribution, provided attribution is given. This means a startup could take Amália and fine-tune it for legal document analysis, or a municipality could adapt it for a local tourism chatbot. The community is encouraged to contribute back improvements, creating a virtuous cycle of shared development.

Portugal’s move is a reminder that in the race for AI supremacy, there is value in going small, local, and transparent. The next two years will determine whether Amália becomes a durable piece of digital infrastructure or a well-documented research project with a beautiful name.


Source: TNW | Artificial-Intelligence News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy