Data and models

Training Data and Model Ownership

Navitra both pre-trains language models from scratch on client data and fine-tunes existing open-weight models. The questions on this page — who owns the weights, where training data goes, and what happens when an engagement ends — are more substantive here than for a typical AI supplier. This page answers them plainly.

Client data trains that client's model only

Data a client provides to Navitra is used only to train, fine-tune, or evaluate that client's model. It is not pooled with data from other clients, it is not used to train a general Navitra model or product, and it is not used in demonstrations, benchmarks, case studies, or examples without that client's explicit written consent.

  • No pooling: each engagement uses a fully isolated dataset. One client's data does not influence another client's model.
  • No product improvement: Navitra does not retain or reuse client data to improve its own tooling, pipelines, or future model training except where the client has agreed to this in writing, with the specific purpose and scope stated.
  • Derived artefacts follow the same rule: processed datasets, tokenised corpora, evaluation sets, and embeddings created from client data carry the same isolation requirement as the raw data they were derived from.

Any departure from this — for example, a research collaboration where sharing certain findings benefits both parties — must be agreed in writing before that use begins, with both parties' rights and obligations stated explicitly.

Who owns the model weights

The trained model weights produced under a Navitra engagement belong to the client. Navitra claims no ongoing right or licence to use, copy, or distribute them. The client may modify, retrain, audit, or decommission the weights at any time without reference to Navitra.

Navitra retains ownership of its training infrastructure, software, tooling, pipelines, hyperparameter configurations, and methods. These are not transferred to the client and are not client-specific; they are the proprietary means by which Navitra builds AI systems across all engagements. Clients who require bespoke tooling to be transferred as a deliverable may negotiate this separately in the engagement contract.

For fine-tuned models, the client's ownership of the adapted weights sits alongside the licence conditions of the base model. See the section on base models and licences below.

Base models and open-weight licences

When Navitra fine-tunes an existing open-weight model, the client's right to use the resulting model is bounded by the licence of the base model. Open-weight does not mean licence-free.

  • Commercial use restrictions: some open-weight model licences prohibit or restrict commercial use, either absolutely or above a usage threshold defined by monthly active users, revenue, or organisation size. A model that is free to use for a small business may require a separate commercial licence for a large enterprise.
  • Attribution requirements: some licences require that use of the model or its derivatives be attributed to the original authors or that the licence text accompany any distribution.
  • Industry restrictions: certain licences exclude specific use cases — including military applications, surveillance, or particular regulated sectors. These restrictions carry over to fine-tuned derivatives.
  • Copyleft conditions: some licences require that derivatives be released under the same licence terms. Fine-tuning a model under such a licence may affect the client's ability to keep the resulting model proprietary.

Before beginning any fine-tuning engagement, Navitra reviews the licence of the proposed base model and raises any material restrictions with the client. Where a licence condition would affect the client's intended use, Navitra will identify an alternative base model, seek a commercial licence from the original author, or advise that from-scratch pre-training is required.

Base model licence review is conducted per engagement and documented in the engagement scope. Clients are informed of any material licence restrictions — including commercial use limits, attribution requirements, or copyleft conditions — before training begins.

Pre-training data for from-scratch models

The source of pre-training data varies by engagement. In all cases, the data sourcing approach is agreed with the client and documented in the engagement scope before training begins. Three configurations are possible:

  • Client-provided data: where the client provides all pre-training data, the client confirms that they hold the rights necessary to use that data for model training. Where the data contains personal data, the engagement documents the lawful basis under UK GDPR.
  • Navitra-curated data: where Navitra sources or curates pre-training data, Navitra confirms the provenance and licence of each source and confirms that its use for model training is consistent with those licences and, where personal data is involved, with applicable data protection law.
  • Mixed: where training data combines client-provided and third-party sources, each component is documented and assessed separately under the applicable scenario above.

The specific data sources for each engagement are documented in the engagement scope and are available for review by the client.

Where training happens

Training runs on cloud compute infrastructure provisioned from Google Cloud Platform, Microsoft Azure, Amazon Web Services, or Hetzner, according to the compute requirements of the engagement and the client's region preference. Where a client requires training to run inside their own infrastructure — on-premises or in the client's own cloud account — this can be agreed as part of the engagement scope.

  • Data residency: by default, training data is transferred to Navitra's managed cloud compute environment for the duration of the training run. The compute region is confirmed with the client before data transfer. Access is restricted to the named Navitra team members for that engagement. Client data is deleted from Navitra's infrastructure at the end of the engagement as described in the Retention and Deletion section below.
  • Hosted compute: training infrastructure is provisioned on cloud platforms with data processing agreements in place. Provider personnel do not have access to client data in the ordinary course of operations.
  • Cross-border transfers: where training data containing personal data is transferred outside the UK, transfer is made under a UK International Data Transfer Agreement (IDTA) or equivalent mechanism. The destination country is confirmed with the client before transfer.

Retention and deletion

A training engagement produces artefacts beyond the final model weights. Each of the following is subject to the same data handling obligations as the training data it was created from:

  • Training checkpoints: intermediate model snapshots saved during training runs. These contain the model's weights at a given point and, implicitly, the influence of all training data seen up to that point.
  • Embeddings: vector representations derived from client documents or data that may be used for retrieval, fine-tuning, or evaluation.
  • Evaluation sets: labelled examples, test prompts, and expected outputs derived from or representative of client data.
  • Training logs: records of training runs including loss curves, configuration, and any outputs sampled during evaluation.
  • Processed datasets: tokenised or otherwise transformed versions of the raw training data.

Default retention period for all of the above: retained for the duration of the engagement and for 90 days following written confirmation of completion, then permanently deleted unless the client requests otherwise in writing before that date. Deletion is initiated by the Director. A deletion confirmation record is retained. Clients may request written confirmation of deletion.

Erasure and trained models. A trained neural network cannot straightforwardly have individual records removed from it. A database row can be deleted; a trained model cannot. The influence of a specific training example is distributed across the model's weights in a way that cannot be identified or reversed without retraining the model from a dataset that excludes that example. This means that if a client is subject to data subject erasure requests under UK GDPR Article 17, and if the scope of those requests could extend to the trained model itself, this must be planned before training starts — not addressed retrospectively. The data at issue must be identified, its role in training must be understood, and the contract must specify what Navitra will do if an erasure request arrives that the parties agree extends to the model. Current technical methods do not reliably guarantee erasure from a trained model short of retraining. Navitra will advise clients on this risk at the outset of any engagement involving personal data.

Confidentiality

Client data, model weights, evaluation results, and all artefacts produced during a Navitra engagement are treated as confidential. This applies from the point of receipt and continues for the term of the engagement and for five years following its conclusion, unless a longer period is specified in the engagement contract.

  • Access: access to client data and model artefacts is restricted to named Navitra staff and contractors who need it for the specific engagement. Access is removed when it is no longer required or when the person leaves the engagement.
  • Non-disclosure: all staff and contractors with access to client data sign a confidentiality agreement before the engagement begins.
  • No publication: Navitra will not publish, present, or reference the client's model, training data, or results — including in anonymised or aggregated form — without the client's written consent.

The technical controls that protect client data and model artefacts during an engagement are described in our Information Security statement.

Subcontractors

The following categories of subcontractor may access client data or model artefacts during an engagement:

  • Compute providers: training infrastructure is provisioned from Google Cloud Platform, Microsoft Azure, Amazon Web Services, and Hetzner, as required per engagement. Data processing agreements are in place with each provider. Provider personnel do not have access to client data in the ordinary course of operations. The provider and region used for each engagement are confirmed with the client before data transfer.
  • Annotation services: Navitra does not currently use third-party data annotation or labelling services. If this changes for a specific engagement, the client will be notified before access begins and a data processing agreement will be in place.
  • Client notification: subcontractors with access to client data or model artefacts are identified in the engagement contract before work begins. Clients are notified before any change to the subcontractor arrangements for their engagement.

A list of third-party processors used by Navitra more broadly — including those outside training engagements — is maintained in our Privacy Notice.

Who owns model outputs

The client owns the outputs their deployed model produces in production use, to the extent that ownership can be established. Outputs produced by Navitra during evaluation, testing, or demonstration are the client's to use for the purposes of the engagement.

Two legal uncertainties are worth noting explicitly:

  • Copyright in AI-generated content: under current UK law, AI-generated works — content produced autonomously by a computer without a human author — may attract copyright protection via the "computer-generated work" provision in the Copyright, Designs and Patents Act 1988, with the protection vesting in the person who made the arrangements for the creation of the work. The legal position in other jurisdictions differs and is evolving. Clients who intend to commercialise AI-generated content should take their own legal advice on the IP status of that content in each relevant jurisdiction.
  • Third-party IP in outputs: a model may reproduce text or other content from its training data in its outputs. If that training data contains third-party copyright material, the reproduction of that material in outputs could constitute infringement. Clients should not assume that outputs are free of third-party IP claims without reviewing the provenance of the training data and taking appropriate legal advice.

Questions about a specific engagement

Questions about how data is handled in a specific Navitra engagement — including requests for documentation, clarification of contractual provisions, or exercise of data subject rights — can be raised via the contact page. For data protection matters, you may also refer to our Privacy Notice.

Version 1.0  ·  Last reviewed 25 September 2026  ·  Owner Jaydev Bhatt, Director