Podcast episode with the Open Terms Archive team

[Commons] Monitor platforms with Open Terms Archive

Open Terms Archive: follow the platforms’ T

Walid : Hello and welcome to Projets Libres, LinuxFr.org’s podcast, where we talk about free software, digital commons and open data. Today, I’m very happy because we’re going to talk about the analysis of the terms and conditions of the major platforms. To do this, we are going to introduce a project called Open Terms Archive. I am very pleased to have two people with me who will talk to us about the Open Terms Archive.

Presentation of the guests

I have Sydney Wheeler, who is responsible for Open Terms Archive partnerships and Matti Schneider, who is director of Open Terms Archive. To start, as usual, I’m going to ask you to introduce yourself, to explain to me who you are and how you discovered the digital commons, maybe free software. And also what is your role in Open Terms Archive as well. Sydney, I’ll let you speak. Would you like to start, please?

Sydney : Yes. So, there is a path, as we like to say, a little atypical compared to the digital commons sector. I started my career more in the world of NGOs. So, I worked with Handicap International, and an association called Helen Keller International, too, for several years. Before setting up my own business, I was still familiar with the digital commons, but a little from afar. And it was rather through Beta.gouv that I really discovered the subject of Open Source software, free software and the digital commons.

So, through a first mission on a project called Zero Vacant Housing, when I started my self-employed activity. So it was in 2022, I think. And after this mission, it opened me up a little to the world of digital development, digital public service and subjects such as Civic Tech, etc. And I discovered Open Times Archive, It must have been in October 2024, when I saw an offer of assignment within the team on the deployment of the product.

Walid: And you, Matti?

Matti : Hello, so Matti Schneider. I’ve been working on the digital commons and free software for some time. I have always sought the meeting between my objective of impact in society, to have more social justice, more links, more connectivity, I would say, between humans. What I have in my hand is the ability to build software and lead teams. So, I started working with the digital commons even before the term became widespread, creating a first social impact startup, with the goal of reducing the carbon footprint of university campuses.

It was based on something that was emerging a bit at the time, which was OpenStreetMap and the creation of Wikis around it: assembling several bricks which, today, seem quite logical to me to rely on digital commons for this kind of objective. But it was a frankly quite burgeoning approach at the time. And I couldn’t say that it was totally thought out, even on the team’s side.

But in the end, it is this approach that I have retained in all the positions I have been able to take since, in particular by working within the French State and other States thereafter. Sydney mentioned in particular Beta.gouv, the French state-owned startup incubator, which I co-founded in 2014 and within which we have been able to build several free software, some of which have emerged and moved towards digital commons.

And in particular, I wrote a white paper on the digital commons in 2017, at the end of my work within the DINUM [Interministerial Digital Directorate], and which was able to give a bit of an axis of reading, of the possibilities of creating digital commons within the public authorities and in conjunction with the public authorities.

I then led several missions with different governments around the world, and I returned to France in 2018-2019. I worked on information manipulation issues within the Ministry of Foreign Affairs. And after a year and a half of trying in all directions, to see what could be done, and in particular with a lot of hope about what was possible to accomplish with the world’s second diplomatic power, we were still forced to conclude that, in the end, no state, no matter how powerful, could influence the digital ecosystem on its own. especially that of Big Tech.

So I went back to what I know how to do, which is to create digital resources with open governance to create broad coalitions. And that’s how Open Terms Archive was born.

What we founded at that time was indeed a technical tool, I imagine that we will describe it a little more, but also and above all, the ability to federate actors around a digital resource. Today, indeed, I’m the director of the Open Terms Archive, because it’s convenient to have a business card where there’s a word like director written on it.

We have a lot of partners for whom it’s important to have something that looks like a somewhat classic structure. But let’s be honest, when you’re a very small team, everyone wears all the hats. What matters is what you do and not really the title, except when you are looking to have institutional partnerships.

What is Open Terms Archive and Its Genesis

Walid : We’ll come back to the genesis a little later. First, I would like you to explain to the listeners what the Open Terms Archive is, what is this project and what it is used for and what are its main features?

Sydney : Open Terms Archive is a digital commons that publicly records every version of the digital platforms’ contractual documents and notifies of changes made to those documents. This is the somewhat technical presentation, but in a slightly simpler language, OpenTerms Archive will allow you to track all the changes that are made to documents such as the terms and conditions of use [TOU], privacy policies, etc., to have notifications about these changes.

There is the possibility, therefore, to look at the changes in the form of a diff. So you can see the original text on one side and the new version of the document on the other. This avoids the need to read the entire document to understand the change made,

In particular, we will save each version, so we can look in the history, what was the state of this document on such and such a date, and possibly compare it to today’s version.

Walid : Matti, do you want to add something on the subject or not?

Matti : No, I think Sydney explained it very well. Afterwards, presented like this, it always gives hope and the impression that we will be able to use it individually, and it’s true.

I think there is a potential interest in accessing … These are huge documents that we have not read and we have ticked the box by saying that we have read them. But hey, let’s be honest, having a version recorded in another place is not going to make you want to read it more. So, by creating this tool and this database, and this entire ecosystem, in the end, we are still aware that it is not the collection of documents as such that will allow individuals to rebalance the balance of power that exists today.

We are always subject to these rules. Making them visible does not exempt us from this, so it is important to understand that the intention behind OpenTerms Archive is above all to equip players who have the ability to influence the major platforms.

When we set up Open Terms Archive, I gathered my team at the time: we made a list of the players who are able to win victories against Big Tech from time to time. We concluded that there were 4 main types of actors. There were the regulators who can impose fines that can be counted in the millions or billions of euros. And there, indeed, it changes behavior. That’s how we managed to have a choice of browsers in operating systems, for example.

Another actor is the legislator, because the GDPR is something we like to reduce to cookie walls, consent walls, etc. But it is also a way of reducing the scope of this text, which is very, very important. He has changed a lot of behavior. I think that we must stay away from this propaganda that we would like to trivialize it. And to believe that it’s just an extra annoyance to tick off, yes or no.

So, the legislator takes a long time, it takes years and years to come up with a text, but when one comes out, it changes something. Why? Because there is someone who is likely to go to jail and that makes the big companies think a little more.

Another player is also consumer protection associations. And in fact, they are the ones themselves, but it is above all this ability that they have to activate the judicial system. Again, because there are penalties that are serious, with potentially practices that will be banned, criminal liability, that kind of thing.

And finally, the media. After all, these media, I don’t think that in essence, we need more proof that Big Tech and most of the big monopolistic players are harmful. It’s not a great discovery and to say, “Oh, there, you don’t have what, they’ve once again ripped you off in their generalization conditions.” So that doesn’t change anything.

On the other hand, once in a while, there is still a scandal that is big enough. So that there is a kind of global reaction that will threaten the user base a little bit. Maybe we remember, a few years ago, the data merge between WhatsApp and Facebook, for example, while Facebook, at the time, had promised, swore, spat that they would never do it. Well, in the end, they did it.

And what was supposed to pass like a letter in the post, like yet another change of taking its use, actually led to a form of reaction. This was clearly maintained by the visibility that was brought by the press and the media. Once in a while, there is still visibility that can be brought by these actors.

So, if we make this list, regulators, legislators, consumer protection and the press and media. A few years ago, I went to interview and talk to many of these people. And we concluded that there was a flaw, well, something that they lacked every time, which is that they are generally bad at tech, but they need more or less the same data sources.

Finally, in order to react, they first need to be aware of the problem and they need to be able to objectify it. If you want to file a complaint, you need proof. And so, we found ourselves talking to people who explained to me that they had taken screenshots of the documents as they were at the time in order to be able to file a complaint. In fact, it’s a horror for them. And sometimes, a year and a half later, they find themselves having to go and pick them up.

So, that’s really where it started. It’s how we can strengthen the capacity of the actors who, today, can influence the platforms and what can we pool in this. This is the first step. The next step is finally if we are able to provide them with this data and we gather them around us and then send them this information. We can also direct these reactions and make them common. And so, instead of having the regulator on the one hand, who will complain, and on the other, three months later, the consumer protection society. And then, in the middle, a little bit of the press that seizes on it.

Finally, if we succeed in having a consolidated reaction, we can significantly increase the striking power and have a reaction that is stronger than the sum of the individual actions. That’s what we wanted to build with Open Terms Archive. So yes, it can be used by the end user who will be able to find his conditions of use. But let’s be honest, the objective is above all to equip those who, today, are already able to exert pressure

Walid : On that point, what I’d like to know is when you say we talked to a lot of people and everything, it was rather in the French ecosystem where you were at the European level, well basically with whom you are talking, which makes you come to this conclusion that we need a technical tool to monitor all these changes?

Matti : This research is being done as part of an incubation within the French Ministry of Foreign Affairs, so at that time, I was the director of innovation for the Ambassador for Digital Affairs. So, in the end, the exchanges are a little French, but a lot European and even international. It’s very broad.

I wouldn’t say that the conclusion is that you need a tool. The conclusion is that something must be done that is beyond the capacity of a government, alone, of a state, alone. And what I know how to do, what I can bring to the table, is the construction of a digital tool. But I mean, someone who would have had a different skill set would have done something else.

This is something quite interesting, which is not the main focus here. But it is to have had the opportunity to do digital diplomacy. The Quai d’Orsay, without digital skills, would probably have come to the conclusion that we are going to create an international forum for discussion on how to do it. There are plenty of regulation and exchange systems to put more constraints. It also works, but it’s another angle of action that we have chosen, which is to mobilize technology.

But as I said, there is no fantasy that raw data will change anything. On the other hand, data as a means of activating an ecosystem and as a means of federation, to consolidate influence, is indeed new. And that, in the end, is the intention of digital diplomacy.

So there you have it, it was really global. And in fact, in our case studies, we see it. We had some very interesting cases. We were directly solicited by the American legislator, and that, for us, was one of the main marks of impact, because as long as it is primarily American companies, When we want to influence them or limit their impact, we can try to issue legal obligations on our soil, but it’s always quite secondary.

We can see that recently, Elon Musk was summoned to explain to us how exactly, he justified the idea of freely making robots available that can undress people without their consent. He just didn’t deny coming to Paris. So, indeed, if we are able to feed the debate and perhaps even encourage the American legislator to increase protections, we have a very important and global effect.

Walid : That leads me to the question of which country Open Terms Archive is most active in. Sydney, do we have large areas? I understand that there is Europe because a priori, that’s where it started. And there, we were talking about the United States. What is known about the places where Open Terms Archive is used in the world?

Places where Open Terms Archive is used around the world

Sydney : To explain the organization of the Open Terms Archive: the documents that are tracked are organized by collection. Each collection is managed and maintained by a partner.

And we have federated collections, so that means that they meet a certain number of criteria that allow us to publish them on our site, to have visibility on what is being done with these collections, who maintains them, etc.

On the other hand, since it is free software, there could be other uses that we are not necessarily aware of. Within the active community of federated collections, Open Terms Archives is used mainly in France and Germany. We also have a collection that is maintained by a group of volunteers in Kenya. in Switzerland and Sweden.

So, Switzerland and Sweden too, they are volunteers who have created and maintain collections. So these are the main locations of contributors. On the other hand, we have documents that are monitored in other jurisdictions, where we do not necessarily have contributors on site, but where we have collections that follow the documents, particularly throughout the United States, Canada and Ireland. Who represents the court? Europe, Australia and the United Kingdom.

Walid : Can you give an example of what a collection is?

Sydney : yes, a collection, It’s going to group a set number of services. It’s the platforms, Instagram, Meta, X, for example, are services, and contractual documents. To create a collection, you need to identify its parameters beforehand.

There is a whole phase of construction of the collection. We will identify which services we want to follow. Often, it’s by theme. We have collections that follow, for example, providers of generative AI models.

We have collections that will follow the VELOPS, the Very Large Online Platforms, the large digital platforms. Or it can be in a jurisdiction with fewer specificities on the theme. Most of the time, the collections are still thematic overall. And then, we’re also going to define what type of document we want to track. So, a department has a whole collection of contract documents. And so, depending on what we want to observe, If we do research, for example, we will specifically select what type of document we want to follow in these different services.

Matti : Another important point, in addition to what Sydney said about the collection, is that they are aimed at a specific jurisdiction and often language. What we discovered is that in the end, a service can offer different contractual conditions from one country to another.

This, generally, as we can see, can change, but they will not always indicate it. That is to say, depending on the geographical location of the computer that will consult the page, we will be offered different content. This, even though the service will claim that it is the same.

A classic example that we discovered is the case of Meta, which claims to have a set of conditions for North America, a set of conditions for the European Union zone, and a set for the rest of the world. So, indeed, according to their site, you can select between the three and browse, then see it in many different languages. Except that what we see with Open Terms Archive is that the collection, if we are really present on computers that are physically located in other countries, we have small variations. And that, to detect without a tool like Open Terms, Archive, it’s absolutely impossible, because it’s a sentence that will change in the middle of the text.

So, for example, discovering that in an African country, we will have Facebook which will allow itself to give the names of people who give erroneous information about the result of the elections, for example. These kinds of things that do raise huge questions, but which, from an external point of view, if we just browse through the options offered to us by Meta, do not exist. And so, you really have to put these probes in different places.

So, that’s why a collection, as Sydney indicated, is a theme that will bring together actors, services in a certain industry. It is a set of types of contractual documents. For example, if we focused on privacy statements, it’s not quite the same as if we wanted to focus on general terms and conditions, because we may be interested in different subjects. But it is also a specific jurisdiction.

So, these are the general terms and conditions of sale of the main marketplaces in France, or in Germany, or in the United States. And it’s not necessarily the same documents, but it’s going to be the same services and the same types of documents.

Examples of things that are discovered with Open Terms Archive

Walid : Yes, on that, I would put two conferences that I used to prepare this intervention. In one of the conferences, Sydney, I think it was you who explained, I don’t remember for which service, in South America, they had cut the thing with, two countries that had jurisdictions that were a little more advanced than the others in terms of data protection. And so, in fact, there were general conditions that were different depending on the progress of the legislator on this.

And what I wanted to know was if you could give listeners concrete examples of what we’re seeing with the Open Terms Archive, in addition to what you’ve already cited, both of you.

Sydney : Yes, indeed, this example is not an observation we made at Open Terms Archive because we don’t have a collection in South America yet. We would like to be able to test. These are partners with whom we have exchanged, who have explained this difference to us. Indeed, Mexico and Brazil are two countries with greater regulation and influence on platforms. And so, the platforms are more likely to submit to these jurisdictions.

But the platforms will draft contractual documents that meet these regulations and will apply them in all Spanish-speaking countries. For example, the rules are based on Mexican regulations. And there is also Brazil, which influences the implementation of the rules by the platforms. But that’s something we learned in exchanges with other players in South America. But we haven’t had the chance to validate the hypothesis with Open Terms Archive yet.

Walid : Matti?

Matti : To grab the pole you’re giving us, thank you, I tell listeners that we have quite a few examples of case studies on our site, especially in the memo section, where you’ll find a number of analyses that are done on the basis of the raw data we produce.

Because obviously, just rereading diffs, changes, it’s not the most exciting. So, in some places, we did it for you. Maybe to give two or three examples: Proton Mail, for example, which is a company that says it provides an anonymous and secure service, which has had a little bit of controversy lately, and in fact quite familiar with controversies.

In particular, one of the first to appear in France in 2021, with the provision of, and finally the transmission of, data concerning climate activists to the French police, who were very surprised to discover that in the end, it was information from Proton. When the scandal broke, Proton defended himself by saying “But at the same time, we can’t do anything about it, it’s Swiss law that imposed it on us, you understand that even if we’re based in Switzerland, if we have a legal order, we’ll still have to follow it, and besides, it’s indicated in our T”

Yes, it was true, except that what we could see with Open Terms Archive is that it had been added to the Terms of Use 12 to 24 hours before the press release, and that this was not the case at the time people had signed. So this,

It’s an example where, in hindsight, we can still rebalance the message a little bit and potentially, then, in this case, there was no complaint, so it wasn’t used for a legal process, but it’s an example.

An example of use in rapid reaction, this was the case during a collaboration we had with UFC-Que Choisir, so the largest consumer protection association in France, which made it possible to have, for example, the director of legal affairs of the UFC, who sends a small email to a large French online marketplace platform, when it decides to change, without warning or indicating it, its delivery conditions.

The detection is done by Open Terms Archive. Because if there wasn’t this tool, in the end, someone would have to reload the page every two days and reread the entire 60 pages. So, no one is going to do it. Thanks to Open Terms Archive, we reverse the system.

The UFC Que Choisir receives a notification that there has been such and such a change, it may be worth going to see. We will see and we see that there are just the three lines that have changed, that are highlighted. And so, the director of legal affairs can send an email saying “I’m very surprised by this change, we’re not sure if it’s legal, could you explain it to us?” As a result, within 48 hours, finally, there is a correction, without having an answer that is given. But we register that the change is cancelled. So that’s another use case.

And you can see, again, that I fall back on something where the average user, He can’t do much about it. But when you are in partnership with, in the first case, as I indicated for Proton, with players in the press or here, with players in consumer protection, you can have a fairly direct effect. But afterwards, it can also be used effectively for individual or organizational behavioral adjustments.

For example, we detect that Mistral, the nugget of French AI, made a change last year, which contains two large blocks. The first is where it updates the list of its data processors by indicating that it will also have Google, based in the United States. And so, for the first time, there is a transfer of data outside the European Union. And in the same change, it also removes their commitment to notify you if it changes the list of data processors. So, until then, you thought that he would warn you if this is the case. But in fact, they are exempting themselves from the obligation and at the same time, they are doing so.

And they do so because it is easy to understand why they are exempting themselves at that time: for the first time, they are adding treatment outside the European Union. So, it’s not very fair, let’s say. So, we are happy to be able to detect this kind of thing. These are very concrete examples of what can be done with it.

Walid : I would put in the transcript a link to a video of a person whose name I have forgotten, which is a hearing at the National Assembly in the commission of inquiry into the vulnerabilities of the systematic system in the digital sector. A person from Brazil who explains very well how they managed to get X banned. etc. It’s very interesting, we talk a little about the subject, in a different way, but it’s really very interesting to listen to on the subject. Sydney, I’ll give you the floor again. Maybe there are other examples or other things that you wanted to add.

Sydney : I also wanted to give another counter-example of how it is ultimately used by end users. This is an example that I spoke about at the event on January 20. Last year, there was a trend of platforms banning political ads.

And we saw that it was put in place in the months before the entry into force of a new European regulation related to the transparency and targeting of political advertising, known as the TTPA. Open Terms Archive has therefore published a small report on these behaviors observed by platforms, including LinkedIn. And one of our contributors was subsequently able to report political ads she saw in her LinkedIn feed, after this change, of LinkedIn’s policies.

It managed to have 25 political ads removed, which, as a result, were now banned, according to the new version of LinkedIn’s contract. So, this is an example of a reaction by an end user thanks to an observation we were able to make with the Open Terms Archive. But it’s also an observation that, for us, was very interesting, also because this European regulation does not prohibit political advertising on platforms. And so, we saw an example where the platforms, we imagine, estimated the cost of really complying with the regulations, of being higher than the cost of totally banning these political ads, well this type of content on their site.

And as a result, it’s a rather interesting example of the behavior of platforms, when we put the regulations and the impact that it can have on the behavior of the platforms, we see that there is an impact that was not necessarily the one provided for by this regulation.

Walid : Matti, did you want to add something?

Matti : For me, this idea of impact assessment, which is made possible by observation, is very important. In the context of regulations, there is often this obligation, or at least this expectation, which is a form of objectification, of the effect. And typically, here, in the context of transparency obligations on political ads, what is required by European regulations is: “if you publish political ads, you have to know exactly who paid, how many people it was made visible, who asked for it, etc”.

Indeed, things that seem quite basic to me. It’s important to know if it’s an oil oligarch who pays for something or if it’s actually someone who lives in my country. So it’s pretty defensible.

And if we just look at the direct effects and do a classic impact study, we’re going to say that there aren’t many ads that today meet these transparency criteria, so it’s useless. And finally, thanks to the Open Terms Archive, what we see is that if there is an effect: it’s just that there are far fewer political ads that are available and published, which, personally, I am very pleased about, since having studied the possibilities of manipulating information, it is a very important vector,

And that, in addition, in France, for example, it is simply illegal in many cases. So, I’m delighted that this market is closed, but I don’t think it would have been part of a traditional impact study, since it’s a side effect. I believe that indeed, if we increase the clean-up costs for polluters enough, and the result is that there is no more pollution, this has a very positive effect and should be included in the impact study, rather than just measuring how many times polluters have paid for suffocation.

Walid : I would like us to come for a second on something I didn’t ask for. Earlier, you talked about digital diplomacy. We have talked about the genesis of the Open Terms Archive. It is incubated within the Ministry of Foreign Affairs. What is its current legal form? Do we consider it to be a public service?

Sydney : Indeed, Open Terms Archive was incubated within the French Ministry of Foreign Affairs until January 2026. So, in January of this year, we finished our incubation period and we started from this institutional framework. And today, we are working on this question of legal structuring. What form will the digital commons take? What form will the services we offer to partners take? And as a result, it’s a big topic today within the team. And so, for the moment, we are working in an informal collective. But in the coming months, we will have a more defined structure, formally, in the form of an association certainly. And then a group.

The Open Terms Archive Team

Walid : Matti I’m going to give you the floor, but just before that, how many people is the team? At the moment, how many people are working around this?

Sydney : Today, the team is made up of Matti, myself, and two developers, Clément Biron and Nicolas Dupont. We also have a community facilitator called Clifford Ouma, who is our community facilitator in Kenya. Over the past year, we have also worked with a team of analysts on a project in partnership with the organization Reset on the monitoring of conditions in English-speaking jurisdictions. And then, we have a whole community of contributors who are made up of our partners and who keep the common alive.

Walid : We’ll come back to the community part later. I have a lot of questions about that. I’ll give you the floor, Matti, on the question of structuring.

Matti : I just wanted to answer your question really explicitly. No, Open Terms Archive is not a public service. For me, it’s important to clarify, indeed, as Sydney indicated, the incubation period has come to an end. And I would say, even earlier, eventually, The intention has always been to build a digital commons. And so, the incubation mechanism within a ministry does not create a framework for a public service mission.

So, no, it’s not a public service. This also means that there are not necessarily the obligations and rights and duties that come with such a system. Today, Open Terms Archive is out-of-state. Even if there are state actors who continue to use it and we hope to continue to increase this number of partnerships and uses, it is not a public service.

Walid : OK, thank you. For listeners, who want to know more, we did two more interviews on state startups. I invite you to listen to the episode on the National Access Point for Transport Data and on the National Building Repository. I close the parenthesis.

So wait. Earlier, I come back to this, you talked about digital diplomacy, and we talked about the fact that the American government relies on the Open Terms Archive. How does the U.S. government hear about the Open Terms Archive? Were you the ones who somehow had contacts and introduced them to them? Is it at the level of French diplomacy that this is put forward? I’m interested in understanding how they come to hear about this French initiative.

Matti : Okay, it’s not the U.S. government. It is a parliamentarian, well, a parliamentarian, in this case, as part of a multi-partisan initiative, well bipartisan in the United States, since multi-equal two there, which is called the TLDR Act, a very nice acronym for “Too Long, Didn’t Read”. In this case, the objective is to make it mandatory for online service providers to provide a clear and human- and machine-readable version of their terms and conditions of use.

And so, they hear about us through our partners, and they approach us more from a technical angle. And that’s where, for me, it’s one of the victories of digital diplomacy, because, precisely, no, it’s not a classic institutional approach. Because, frankly, an actor, an emanation of a foreign state could never go and lobby an American congressman. And I mean, in a way, so much the better, it would be borderline interfering.

On the other hand, that the French state incubates the creation of expertise while letting a community take hold of it, without giving direct orders, but by saying: “there you go, we have the impression that this subject is important and as a public actor, we have an interest in such a community emerging, so we support it financially and through visibility, etc”, creates an area of expertise and knowledge that is not directly under the control of the State, but that can be requested by third parties.

And so, we will strengthen the capacity of actors who are aligned with this intention of change in other places. In this case, typically, for the American parliamentarians, they would have presented their initiative in any case, and that is when there is no interference. On the other hand, we will strengthen the credibility of their proposal, since when they present the bill to Congress, it is not an empty idea. We are able to say yes, it is possible to collect, compare, simplify the general terms and conditions of use.

By the way: there you go, a database already exists, that’s the ecosystem, that’s what it has allowed when the general terms and conditions of use are made more readable and analyzed. So, it is in this place, in the end, that there is indeed a victory for digital diplomacy, but it is really a form of diplomacy in its own right, and not just a classic diplomacy that would be equipped by digital technology.

Analysis of the T of different parts of the world

Walid : I come back once again to the analysis of the general terms and conditions of use. Does having analyses on the same general conditions, but in different areas, teach you things by looking at them by comparing them. Does this tell you anything about how these platforms work and the differences they can bring from one territory to another? Sydney?

Sydney : It was a project that we have been carrying out since September 2025. And it’s something we’ve wanted to do for a long time so we were delighted to have the opportunity to set it up. And it must be said that we were at the beginning of the project, we weren’t sure that in the six months of the project that we had, of observation, sorry, that we were really going to get relevant observations.

And in fact, it is. Even in a short period of time, we still saw some pretty interesting things.

So we have seen, for example, the reaction of platforms to the implementation of fairly strict regulations on Australian jurisdiction in relation to child protection. This subject is now a common topic in the European debate, but we have seen real actions taken on the Australian jurisdiction before it is the trend we see today in Europe. So, already, we can see the differences in the subjects that are being put on the table in different jurisdictions. For us, it’s super interesting. So, having the different jurisdictions allows us to see that.

We have also seen differences in the details given in terms of the general conditions of use between these jurisdictions and according to the regulations. For example, in the United States, there is a fairly weak regulation. The example I have in mind is the conditions, the privacy policies. As a result, what is done with user data is not very detailed in the American version of this document.

On the other hand, in the European version, we will already find a much higher level of detail. And we could see that in the United Kingdom, where there are additional specifications, we have even more information in these documents. So, this one is also an example. The importance of this multi-jurisdictional comparison, which allows us to understand, or, in any case, to observe, the impact on the behaviour of platforms that regulation can have.

Walid : If I come back for five minutes to finish on the subject of governance, what you were saying earlier is that these collections are managed in a decentralized way, but you, as the Open Terms Archive team, are you responsible for certain collections. And you manage them? Or is it in any case communities that will take it up according to their interests?

Sydney : Indeed, when a collection is managed by a partner and the collection is federated, so visible on our site, it means that the partner has set up a governance of the collection that guarantees the quality of the data, that it will last over time, that there is really a system in place to continue the maintenance of this collection. We, as a team, have been able to start collections through personal initiatives, for example, or we can also have partnerships, which finance the creation and monitoring and maintenance of the collections.

In particular, this whole six-month phase of observation of documents in five different jurisdictions, these are collections that we have created and that we have maintained within the Open Terms Archive team thanks to a partnership with an NGO that has financed this work. Today, these collections are not federated on the site because we do not know today what our ability will be to maintain them in the long term. So, the databases are available on GitHub, but on the Open Terms Archive showcase site, they are not found for the moment.

Walid : How can individuals contribute to collections? Can they, for example, add new services for which we want to follow the general conditions? What can they follow? What can an individual do?

Sydney : Yes, so. We have two contributory collections. There’s a collection called Contrib. Anyone can propose new terms to follow within this collection. And to do this, we have a small contribution tool that is a bit homemade for the moment, but we would like to be able to improve it in the future. But in any case, it allows you to copy and paste a URL address and therefore propose new terms to follow. Afterwards, if you have a bit of technical skills, you can write your statement directly and submit it via GitHub.

And we also have a GenIA Contrib collection, which is a contributory collection of pursuing the terms of generative AI services. So that’s the same concept, in particular, it can offer a desired service. And this collection is maintained by a private individual, a volunteer contributor to the project.

Matti : Generally speaking, this question of “what collection can I contribute to?”, as Sydney indicated, insofar as it’s decentralized, in the end, it depends on the rules of contribution of each collection. For example, we have a partnership with a consortium of German universities. In this place, the data is used for research purposes and they are not really open to third-party contribution, because it is maintenance work.

I can find that it would be very important to follow Darty’s terms of use. It’s not really their problem at the time. So… They are not open to third-party contribution, but the data is public. And conversely, indeed, as Sydney indicated, there are collections that are specifically designed to be contributory, but that also means that maintenance is rather voluntary and therefore uncertain. So, we’re going to have a contribution and then it can take between a day and 20 days before someone takes care of retrieving and validating them.

So we finally focused on tracking additional documents, but there are many other ways to contribute to the Open Terms Archive. First of all, give meaning to the data that is already present, since collecting data is good, but as we have indicated, it is only the first step. If the objective is to have influence, to allow reaction, to generate diff, it is not quite enough. That’s better, but for there to be a reaction, it has to be the resulting analysis that is put in front of the eyes of the right people.

Finally, even without technical skills, you can subscribe to the news feed by an RSS feed or by other means, [receive] changes that are detected by Open Terms Archive collections, and write a small editorial, a small memo. We have writing guides on this.

For example, if you work in an organization that does advocacy, it can be quite relevant to subscribe to the changes and use them to inform internally. In the end, there too, it is a form of use. And then disseminating, helping to support the visibility of the product, obviously, is always useful, strengthening the campaigns that are made.

Sydney gave the example of an important contributor, Marie-Pierre Vidonne, who has on many occasions contributed subjects, far beyond the technical. The act of activating what you detect by going to LinkedIn to tell them. “But wait, I don’t understand, you changed your terms and conditions of use yesterday, that political ads are forbidden, yet I still see them, here is a selection.”, and that it enforces the law, it has just as much, if not more, value than detecting change.

In the first place, the contribution to improving the quality of the information space involves many other means than simply the collection of additional data. I would say, in a way, technically, we would be able to massively increase the number of documents monitored, whether through contributions, or through automated tools. There is a reason why, strategically, we have chosen not to do this en masse. I’m not interested in claiming 40,000 sites followed. Yes, 40,000 sites, followed, But where no one cares, it’s useless.

We prefer to have only 1,200, but where there are really people who are interested. So, the contribution is indeed done by collecting data on subjects that interest you, but it is also done by analyzing existing data and by activating the rights conferred on you when you change less detected.

Walid : Before coming back and talking again about the community, about the way we contribute, there is a subject that I was not necessarily aware of and that I discovered by listening to the conferences, and that is data normalization, that is to say that we get a lot of data from many different sites, it’s written in the same different way, etc. What is the process of data normalization to arrive at something that makes it more or less easy to analyze data from different sources? Where do we go, in fact? I don’t know which of the two wants to explain.

Sydney : So yes, indeed, we’re going to take back the documents that have their original form, which is very variable. It can be in PDF, it can be in text directly on a page of a website. And we’re going to standardize this form by putting everything in Markdown format, which makes the documents readable from each other. We work, as Matti said, with a lot of researchers, who will need to read a lot of these documents in a row.

We learned that there is a visual fatigue that sets in, as a result, switching between different fonts, different document formats. And so, this Markdown format, which is also going to include multi-page consolidation of the document, It allows them to avoid this visual fatigue. To have more readability on the documents. We will also assign line numbers on the documents, which also makes it easier to read the diffs and understand exactly where. In one document, a change has been made.

Walid : Is the process the same from one collection to another? Depending on the collections, are there small differences that need to be adapted?

Sydney : It’s the same process from one collection to the next.

Walid : By the way, the fact of regularly going to retrieve this data, is it done centrally or from what I understand, potentially, it is done on computers that are in certain different places?

Sydney : Yes, each collection is hosted on a dedicated server. As Matti explained earlier, we have a challenge to have a server that is located in the same jurisdiction as the collection, because the services will often automatically refer to the documents in the jurisdiction of the server that will retrieve them. And as a result, each collection is hosted by a dedicated server that is located in the same jurisdiction as the documents it will retrieve.

Walid: Matti?

Matti : yes, just to try to put it in simpler words, because you must be dumber, sometimes, when you want to follow the documents in German, in Germany, if you do it from Spain, and even if you tell the service that I want to have the German version, you’re not sure they’re going to do it. They can never be trusted. That’s one of our design principles, is that we never trust platforms. We have had too many examples of cases where they lie, where there are abuses.

So, indeed, a server, that’s what Sydney said, a server, a computer, which is geographically located in the targeted area. But it also has another advantage: given the amount of documents we will be tracking, if we had a single centralized computer, we would be blocked everywhere. And so, in the end, we have a lot of computers that collect few documents. So that’s another advantage.

Also, as we have indicated, this is a decentralized initiative, which is very important. If we were the only ones to operate a big server, in the end it would be enough to put pressure on us to stop, to falsify the data. And that’s technically impossible with a lot of small computers that are managed by a lot of different teams. It’s impossible to put pressure on everyone at the same time.

And besides, this is one of the explanations why we are quite comfortable with the fact that there is sometimes redundancy. That is to say, there will be the same document that is followed by different teams in different collections, because a service is both a major social media, but also a velops as designated by the European Commission, but also a major French operator.

In the end, you can have the same documents that are followed in three different collections. That’s fine, because if all three say the same thing, it’s fine. If there is one who starts recording very strange things, we can investigate if it is a technical problem or if it is a corruption problem in a given place. So we consider that, precisely, federation is part of the data security modalities.

Data licenses

Walid : And the data found on GitHub, under what licenses are they? What license did you choose for these documents?

Matti : In the same way, since it depends on the collections, it is each time the data producer who will choose the license he wants to put on the data he produces. In the Open Terms Archive team, we provide the means for collection. But we don’t take responsibility for this collection in a systematic way.

So, each time, the operators will decide on the license. So, we have one that we offer by default, which is effectively an ODC-By license, so Open Database Commons, which allows the reuse, sharing and adaptation of a credit that is given to the maintainer of the collection. But we don’t have any obligation or guarantee that is provided on the licenses. On the other hand, as Sydney indicated, in order to be federated, therefore to be highlighted on the Open Terms Archive site and discoverable automatically via our APIs, it is necessary to meet a certain number of quality criteria, including the fact of publishing the data under a license that allows reuse. It is not our role to interface and make discoverable data that we would not be allowed to use.

Walid : Last question on that, before asking a few questions about the community. The code you use, two questions on that, it is written in what way and the same, what license is it? It’s available on Github too, I guess?

Open Terms Archive Code Technologies and Licensing

Matti : It’s a technical stack that is based on JavaScript, on Node, mainly, which is under the EUPL license, European Union Public License. It’s a Copyleft license that provides safeguards and obligations very similar to those of the GPL, which is a bit more well-known.

The choice of the EUPL is an important political choice, since the GPL family of licenses is based on American law concepts, in particular in the notion of Copyright. For people who have already read it, there is a bit of jargon that does not necessarily make sense in the European context, where the notion of copyright is related, but not quite equivalent to that of copyright.

So, of course, licenses are still contract law. What matters in the end is the commitments that are made, that it is enforceable. And in the end, if I go to see the judge, what will he decide? But we will greatly facilitate the application of the license by choosing a EUPL license, since it is legally enforceable in all the languages of the European Union. The translations are not there to look pretty by saying “there you go, but only the English version is authentic”. No, no. Everything has been negotiated, argued, built with all the Member States, so I can have a Romanian version, and go to Romania to complain. The judge will be able to decide on the basis of what he has provided, and not on a sworn translation, which is uncertain. So it seems useful to us.

And then, obviously, there is an issue, once again, of digital diplomacy, and to show that it is possible to produce software, and, in this case, free software, without relying on American legal layers, but that we are able to create our own operational framework.

People who are able to take advantage of Open Terms Archive data

Walid : Okay, very interesting. I would like to come back for 5 minutes on the community part. We talked about individual contributors, we talked about associations that will use this data. What I would like to understand is in fact the raw data, it is there. What are the profiles of the people who will be able to read this? They are lawyers. What are the profiles of the people who are most likely to understand this data and to get something out of it?

Sydney : That’s a good question because we have several ways of presenting the data we collect with the Open Terms Archive. So, as Matti said, we have different main targets: regulators, legislators, journalists.

We work a lot with researchers and associations, especially consumer protection associations. And the use of data will define in a way how data is consumed. So, for example, journalists will use the memos or reports that are on the publications page. They already have a layer of human analysis done on the data. So, there will already be a first work to popularize what has been observed with the Open Terms Archive, and which will make it easier to reuse the information quickly.

As a result, in the context of publications, for example. Researchers, for example, they will work more directly with Open Terms Archive. So, as Matti said, we have a collection that is maintained by a research institute in Germany, well a university research laboratory. And they will maintain the collection. So they’re going to work from the diffs, so specific changes, but also from the full database.

You can also download the data from the Open Terms Archive to have a much more advanced processing behind it. So the way in which the data produced by the Open Terms Archive is consumed will depend a lot on the skills of the person, but also on the final objective of the use of this data.

We also have, for example, the example of the consumer protection association UFC Que Choisir, with whom we worked to have a collection that would allow them to access the historical versions of documents, because, precisely, they may have complaints or procedures that will concern the document on a specific date. They need access to the version on that date, not the most recent version.

Walid : For example, at UFC Que Choisir, it is lawyers who will look at these documents?

Sydney : Yes.

Walid : Okay. Matti, did you want to intervene?

Matti : Yes, just to say that there are also quite many cases of individuals who will self-organize. The cases that Sydney described are the most important to us in terms of influence. But for the final administrator, there are many other possibilities.

For example, ahead of the early legislative elections that took place in France , we have a small collective of a dozen people that has been created and which, in two weeks, has followed the conditions of use of the main French digital public services. They say to themselves that we need to have a trace, if the far right ever comes to power, of what is done with our personal data and so on. And in the end, it was a very particular emergency: the use case for this data is above all the creation of a history or a collective of users who follow the terms of use of online dating applications. There you have it, everyone, in the end, can build a use case that is adapted to their needs and context.

Funding for Open Terms Archive

Walid : If we talk about financing now. The main funding, at the outset, is the Department of Foreign Affairs, as I understand it. What other funding did you receive? We talked about different programs, partnerships, etc. How is Open Terms Archive currently funded?

Sydney : Today, and even before the end of the incubation period within the ministry, we have different types of funders.

The much-publicized document observation project in five English-speaking jurisdictions was funded by a U.S. NGO, a private association. We also have researchers who maintain their collections, who finance improvements to features that are useful to them. And we can also have contracts with public agencies: for example, we have worked with the digital health delegation, the Ministry of Health, on a collection that tracks the contractual documents for digital services that are referenced in the catalogue of my health space. So, today, we have several types of financing from private and public actors.

Matti : We have also benefited from several NGI funding, so Next Generation Internet, a funding program from the European Commission, distributed mainly by the NLnet Foundation, which has allowed us to have many improvements on the engine itself.

After that, I think about the question of financing. It’s always a bit complex to respond to this aspect in the context of a digital commons, because when you have third-party contributions on the engine, how do you value them? It’s always the question of valuing volunteer work. Should I go to my contributors and ask them how much they are paid by their employer per hour? And how many hours did they spend on this pull request? It’s a bit tricky.

Finally, we are able to measure the funding that has abounded during the time of the team, to talk about it. But all the extra work time, which is provided by third parties, that one, It’s much harder to value. So, indeed, the majority of the funding has clearly come from the French state through the Ministry of Foreign Affairs’ incubation program, but also through mutual funds, for example, as part of the Covid recovery fund. Here, it’s European money that is being distributed by France through France Relance, so I don’t even know who exactly it comes from there.

And indeed, American philanthropy, especially Reset, which Sydney mentioned. Then, we also had a few small grants, for example from Github or the Digital Public Goods Alliance, players in the world of digital commons and Open Source who support us.

But it’s clear that these amounts, in the end, allow us to have some small activities. Indeed, the problem of financing is present and it is all the more important today that we have come out of this incubation phase.

Why? First of all, because this funding from the State, from a public actor, provided us with a form of structural financing, funding that was not associated with the delivery of a particular improvement. What is absolutely exhausting in the context of the production of free software and in the context of digital municipalities in particular, is that we spend our time chasing money. Today, I think that 60% of my time is dedicated to looking for funds and partners.

When, frankly, at the beginning, it wasn’t really that, my thing in life. I want to build software that has a social impact, and today, I spend my time looking for money so that we can do that. It’s a bit of a shame.

And especially today, the funding we have is above all associated, i.e. with the provision of services based on the software. But we are dealing with a form of exploitation of the product and not in its construction and improvement.

We have funding through grants, in particular NGI, which we are extremely happy about today, because it is thanks to this that improvements are made to the software in the first place. But we are still in the project and not in the structural. That is to say, you always have to invent something new to provide, but no one can ever pay for maintenance and annoying stuff. Except that in fact, maintenance and annoying stuff are still the main expense item. And at some point, I’m willing for us to continue to invent new incredible things, but already, if we could have a stable base that continues to run, we’d be very happy.

Challenges for Open Terms Archive

Walid : For listeners who wouldn’t have, then I don’t know how it’s possible now, if you’ve been following for a long time, listen to the episode on NLnet, the NLnet Foundation. I strongly invite you to listen to him, and in particular, precisely, this issue of financing maintenance, since we finance functionalities and not maintenance. And that’s a real problem.

Before we get to the end, I would like you to explain the challenges that are coming? We talked about the challenge of structuring, making a legal structure, etc. What are the challenges that await you in the months and years to come, concretely? Sydney, do you want to start?

Sydney : Yes, that’s right. In the coming months, our main challenges will be the creation of a legal form and the definition of the structure of the Open Terms Archive for the rest of this incubation period within the Ministry.

It’s also going to be financing, as Matti just explained. Today, we have major challenges in terms of the structural financing of the activity. We have a challenge to define our positioning within the ecosystem. Today, we are a data provider. We had a first project on the production of analysis, which is something we did before, but in a much more incidental way and according to the team’s availability.

Whereas here, we experimented, really worked with analysts, with specific skills over a dedicated period of time, and that gave us a lot of ideas for the future. So, we have a lot of work to do to see in what direction we develop?

Final Words

Walid : Before we leave, I’ll leave you a concluding word. What message do you want to convey to the listeners of Projets Libres?

Sydney : I would say maybe since I joined Open Terms Archive, I’ve discovered a lot about the strength of the collective. It’s something I knew from before, from my previous experiences, but in the world of the digital commons and especially issues of platform governance and digital democracy, it’s something very strong.

And so, through Open Terms Archive, I’ve been able to get to know a lot of super interesting collectives, which are really putting together work, which I think, for most of the world, remains a little hidden. The digital world can be scary, it’s a bit technical, but it’s a question of digital rights are really very important, more and more important today. It’s a sector that is really worth watching. I am very happy to be involved in this sector today.

Walid : Matti, I was supposed to let you have the floor for the last word, but it turns out that the recording platform screwed up today and I don’t know why, it’s the first time…

We’re going to stay on the last word of Sydney. It remains for me to thank you both for taking the time to come and present the Open Terms Archive. I’m very happy because, I discovered it a bit by chance, over the course of the conferences, I wasn’t very aware of the subject, I didn’t know it existed.

Around me, when I talked about it, people didn’t know about it either. I was really interested in all of this, so thank you very much for coming. For the listeners, listen, I think you need to talk to those around you, that you will watch, maybe contribute, if you have the time or skills, why not look at this.

Thank you Sydney and Matti thank you very much for your time and then, see you soon.

Sydney : Thank you very much.

Episode production

  • Remote recording on June 15, 2026
  • Basis: Walid Nouh
  • Editing: Walid Nouh
  • Transcript: Walid Nouh

Use of AI

You can consult our AI charter.

Production:

  • Noise reduction in Audacity via OpenVino and the DeepNetFilter2 model
  • Transcription: whisper-medium via mufidiwiwhi locally
  • Transcription improvement: gemma4-26b locally via internal tool (21k tokens output)

Publication:

  • Automatic translation into English: Microsoft Translator with WPML WordPress plugin
  • English translation of social media posts: locally with Jan and Mistral-Small-3.2-24B-Instruct
  • Research of sources around Radxa and an Ubuntu Touch conference: Perplexity

License

This podcast is released under the CC BY-SA 4.0 license or later

, , , ,