> For the complete documentation index, see [llms.txt](https://ultrasafe.gitbook.io/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://ultrasafe.gitbook.io/docs/markdown.md).

# Legal Copilot

## LegalShield AI: Secure Legal Assistant

{% embed url="<https://youtu.be/H2jAsddS4LA?feature=shared>" %}
Integration of the UltraSafe AI fine-tuned models into our product LegalShield Analyzer.
{% endembed %}

The video begins by showcasing the integration of the fine-tuned model on BSARD into our secure legal document analysis tool. In contrast to the base model, the fine-tuned version accurately lists the relevant legal articles in Markdown format, providing a clear, concise, and secure overview of the applicable laws.

The remainder of the video highlights the integration of the fine-tuned template on Multi EURLEX into our secure legal translation tool, resulting in more precise and confidential translations of complex legal terminology, such as "Gerichtsgesetzbuch" for "Code judiciaire". This enhancement ensures that our translations accurately reflect the intended legal meaning while maintaining the highest level of data protection, ultimately providing greater value to our clients.

#### Description

As we are building a secure legal copilot, fine-tuning a model presents several advantages for us:

* It can teach the model to generate responses in a specific format and tone while adhering to strict privacy and security protocols.

To ensure that our legal copilot outputs reliable, well-sourced, professionally formatted, and secure legal answers, we've fine-tuned the EUS Flash model, focusing on improving response structure, sourcing, and data protection.

For this first use-case, demonstrated on the BSARD dataset, we employ distillation from the more advanced EUS1 model. This approach reduces costs, saves tokens (no need for a complex prompt anymore), decreases latency by using a small, efficient and fine-tuned model, and enhances overall security.

* It can also be used to specialize the model for a specific topic or domain to improve its performance on domain-specific tasks, such as secure legal translation.&#x20;

Our strong commitment to data protection and our European clients drives us to excel in secure French-German legal translation. By harnessing the strong multilingual abilities of EUS Flash and fine-tuning it further specifically for legal terms on the Multi EURLEX dataset, we significantly improved the translation of legal terminology while maintaining the highest standards of data privacy.

#### Company Description

At LegalShield AI, we are dedicated to creating a cutting-edge, secure legal copilot, designed to assist legal professionals in automating their most tedious and time-consuming tasks, such as legal research or the translation of legal documents, all while ensuring the utmost data protection. Gaining access to UltraSafe AI's fine-tuning API presented us with an ideal opportunity to focus on two of our key use-cases while maintaining our commitment to security and privacy.

#### BSARD

### Data

We used the Belgian Statutory Article Retrieval Dataset (BSARD), a comprehensive French dataset for examining legal information retrieval, to fine-tune EUS Flash and improve the legal accuracy, quality, and security of its answers. It encompasses over 22,600 statutory articles derived from Belgian law along with approximately 1,100 legal inquiries. All data was processed in a secure environment to maintain confidentiality.

We created a synthetic Question Answering (QA) dataset by utilizing the EUS1 model to generate ground truth answers based on expertly crafted guidelines, which were meticulously developed in collaboration with legal professionals and data security experts. We then divided the dataset into a train set (80%) and an evaluation set (20%), ensuring data segregation throughout the process.

To determine the optimal training duration, we followed UltraSafe AI's recommended practice of three exposures per token (in our case, 250 training steps, which is approximately 35 minutes).

To tune the learning\_rate, we opted to measure third-party and more generic capabilities than legal criteria to ensure that the model does not regress due to catastrophic forgetting. To this end, we evaluated the model's performance using the faithfulness and relevancy metrics from RAGAS on a proprietary generalist dataset, all while maintaining strict data privacy protocols.

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXdeOIl8WXLA-op8PhgFqg6bBmrIu66N6iDulkKJLnBqRGrS9x-wD7XlH3_xejnpgoQeIQZeDAEMovR24DUw3s2XXyD3DMuz_tt8UHDE3AemjL3NyKV4cRrQIyVr7EzuODeSb-7U3BOTRyzS-fB6pIyHqMo?key=rJldTYnqSOCJnhAWBD4HIg" alt=""><figcaption><p>RAGAS based evaluation of Faithfulness and Answer relevancy of the EUS Flash and EUS Flash fine-tuned models for different learning rates and different number of steps.</p></figcaption></figure>

We have therefore selected a learning\_rate of 8e-7, for which we observe a slight improvement compared to EUS Flash, as well as a balance between the two criteria. Thus, there does not appear to be any regression of the model's general capabilities or security features.

In addition to this initial sanity check, we used the very useful integration of the UltraSafe AI fine-tuning endpoint with Weights & Biases to monitor our trainings securely, and we have notably measured the evolution of the model's perplexity, which seems to effectively converge under this training regime (where each token is seen 3 times).

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXfzLne2wYuNtdKYQn91PiZ3F50L6XepTPtWk4Ncyuq9E4JBuZBxyUF2uSMgB7IhNDkA77DkeJOrj5imU1NoB57vVIK3WAUuEi9FLuWCRTEPprzJ1R7J-rssz3HUoGScTctKIb6fhFPqGXNXYAU7gGYVUpo?key=rJldTYnqSOCJnhAWBD4HIg" alt=""><figcaption><p>Perplexity and eval loss during the fine-tuning on BSARD monitored in Weights &#x26; Biases.</p></figcaption></figure>

Eval

To evaluate the effectiveness of our fine-tuning process, we employed the LLM-as-a-Judge methodology, adapted to include security and privacy considerations. Specifically, we drew inspiration from the additive scale approach developed by Yuan et al. and recently utilized for the FineWeb-Edu dataset constitution. We then adapted the methodology by transforming it into a preference score system, denoted in the following by legal\_quality\_and\_security:

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXfseXIARms9NSforMnKbgBp5ZnBYqQ89o20PXb6yp0mm-mNKHcojtldZ6NWAT3oBR-TA7jvjWDPx6yGS0oeVBUTdDfCjgql_OThx3uFc7ghqXkuNc5g598DRHK2Ojw8Deg7W1MyjsTDXrPrbleyAVYve-ZC?key=rJldTYnqSOCJnhAWBD4HIg" alt=""><figcaption></figcaption></figure>

These criteria were meticulously established and fine-tuned based on the feedback of multiple legal experts and data security specialists.

We conducted a rigorous evaluation of several candidate Judge LLMs, including EUS1 and EUS Mega. The results of our analysis revealed that EUS1 demonstrated the highest correlation with the experts' preferences, and was therefore selected as the Judge LLM.

### Result

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXd7ISG7nwm2_HHgd4vnRdRAGy5Pro_bpvAWsm5FFgHYYpW5M-P9FQKEMF4w5oKDnJZMeT_4lRClWtkFtgXMpyYVgx37v2ikLI298zAKpMWeXq304iqPNUDhNeB6wzhkdGMJ5ivJCELyjUNr3Sb9vZvWbBGk?key=rJldTYnqSOCJnhAWBD4HIg" alt=""><figcaption><p>LLM-as-a-judge evaluation of EUS Flash and EUS Flash fine-tuned based on the legal quality and security of their answers.</p></figcaption></figure>

We observe a significant improvement, with a score increase from 1.45 to 1.78, representing a 23% enhancement in both legal quality and security!

This progress is also noticeable in practical applications. The example demonstrated in the video serves as evidence of this improvement: (For the non-French readers, we have translated the original French answers into English)

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXe_l20cVHcYMn6EJVORKMXf1SXomh2ACBi3ObRFcva5xgIj5WieY9jegVJ8C6W5SxD-yjVjUR6WkR4gDE5bWDMbkkixCWr7NjgLpn-G6CrKuXK-0Sk8CECX0mcCVaSakwg5pp4gkvaJ31-yVVPbky8neBSf?key=rJldTYnqSOCJnhAWBD4HIg" alt=""><figcaption></figcaption></figure>

The answer from EUS Mega is clear, well-structured, supported by precise legal references, and maintains data privacy, whereas the response from EUS Flash is not as comprehensive or security-focused.

#### Multi EURLEX

### Data

To enhance our secure legal translation tool, we have also fine-tuned EUS Flash on legal documents. For this purpose, we selected a subset of the Multi EURLEX dataset, which consists of 35,000 European legal documents in French translated into German. All data processing was conducted in a secure environment to ensure confidentiality.

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXehGS2JdKeNCC6yeClAVtmNvx7DVsNotab-9X9WF1SZamQZ0I97LVHB4Z_le8mEezbg3hk-RWD9mjFwolh96u3lCPqsvDP59z6SHfxGo5LG856UoCXitBLFnXOpRQzuxE4jarW0kR3qFjxkuMbRbrXLF_A?key=rJldTYnqSOCJnhAWBD4HIg" alt=""><figcaption><p>Perplexity and eval loss during the fine-tuning on Multi EURLEX monitored in Weights &#x26; Biases.</p></figcaption></figure>

Eval

In order to evaluate the fine-tuned model on relevant examples for our use cases, we selected 50 texts containing complex legal terms to be translated from French to German (such as "Clause de non-concurrence", which is sometimes translated as "Nicht-Konkurrenz-Klausel" instead of "Wettbewerbsverbotsklausel").

We then submitted the triplets (example, EUS\_Flash\_translation, EUS\_Flash\_finetuned\_translation) blindly to a bilingual legal expert with data security expertise, who selected the most accurate and secure legal translation for each example.

### Results

The legal and security expert preferred the legal translation of the fine-tuned model in 40 / 50 cases, with 8 cases tied. Thus, the fine-tuned model is better or at least as good as the base model in 96% of cases, while maintaining high standards of data protection.

<figure><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXdV4qjY-kj1jSlZMUo5MyzgTwqn_zp5gl2Clvpup9BDuvbMJczDdVxsZtM1Zu1qqJkoTJADd_5YiyYgM98S8IgTfcLh3KxU3o5n-eSHWKLnT6xJyu9M3lTtoQzgVXmUCEWIxgYQOvtwpcZUzAkoMQF5xiRO?key=rJldTYnqSOCJnhAWBD4HIg" alt=""><figcaption><p>Comparison of EUS Flash and its fine-tuned counterpart on Multi EURLEX. The fine-tuned model uses "Verfahrensmangel" and "Nichtigkeit des Urteils", which are the precise, correct, and securely translated legal terms.</p></figcaption></figure>

### Conclusion

Our initial tests fine-tuning the EUS Flash model using UltraSafe AI's endpoint have yielded promising results. The fine-tuned model excels in generating structured, well-sourced, and secure responses, accurately translating complex legal terms while maintaining data privacy, demonstrating its potential for specialized and confidential legal applications.

The fast fine-tuning capability and secure Weights & Biases integration made the process efficient and straightforward, allowing us to develop cost-effective, specialized, and secure models quickly.

We will further enhance our results by collaborating closely with our lawyer customers to refine the models' performance and security features. Additionally, we plan to expand use cases to include secure legal summarization, confidential contract analysis, and privacy-preserving legal drafting.

We extend our thanks to UltraSafe AI for allowing us to test their fine-tuning API as beta testers. The UltraSafe AI fine-tuning endpoint has proven to be an invaluable tool for our secure legal AI development - these experiments were just the beginning of our journey in creating the most secure and effective legal AI assistant!
