> For the complete documentation index, see [llms.txt](https://trust.memori.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://trust.memori.ai/en/legal-documents/gpai-general-purpose-ai-models/public-summary-of-training-content.md).

# Public Summary of Training Content

Public summary of training content — MemoriNLP | Memori s.r.l.

{% hint style="info" icon="clipboard" %}
This document is drawn up in accordance with the European Commission Template for the public summary of training content of GPAI models, required by Art. 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). It is the only GPAI document of Memori s.r.l. freely accessible to the public.
{% endhint %}

### What the document contains

* **Provider and model identification:** company name, contact details, name and version of MemoriNLP, date of placing on the EU market
* **Training data modality and size:** type of content used, languages represented and estimated overall volume in tokens
* **Data sources:** publicly available datasets, private third-party datasets, crawled data, user data and synthetic data — with explicit indication for each category
* **Compliance with text and data mining rights:** measures adopted to comply with Art. 4(3) of Directive (EU) 2019/790 and the copyright policy required by the AI Act
* **Removal of illegal content:** approaches and criteria applied in dataset selection and cleaning
* **Linguistic characteristics and structural biases:** distribution of the corpus and mitigation measures adopted

***

### Why this document is public

The AI Act requires providers of GPAI models to make available a summary of the information on training data, for the benefit of **downstream providers**, **deployers** and the interested public. This document does not contain proprietary technical details or commercially sensitive information: for that documentation, see the **AIO (AI Office)** document, available upon formal request.

{% hint style="info" %}
User data remains in all cases the exclusive property of the customer. Memori does not collect or use interaction data to train, modify or update the MemoriNLP model. No personally identifiable data was included in the training corpus.
{% endhint %}

{% file src="/files/kd5OO5j0RDsrngvekya4" %}
