How Should Data Circulate in the AI Era? The Duties of Users and the Rights of Providers
A proposal for sharing distribution load while preserving provenance and usage terms—and returning compensation when appropriate—in an AI-driven data economy.
I want to gather data locally for use with AI.
Having used AI myself, I began to think this way. Instead of searching only when needed, I collect various materials in advance and decide later what to do with them. I find myself wanting to use it that way.
What caught my attention was wondering what would happen if everyone started doing the same thing.
Access for humans reading one by one differs in scale from access for AI gathering vast amounts of data. If many people each delegate collection to their AI and repeatedly retrieve from the same distribution source, servers and networks will bear a heavy load.
Furthermore, if the collected data is processed and those results are then used by other AIs, it becomes difficult to trace where the original information came from.
I believe that in the AI era, we need to consider not only how data is used but also how it is distributed and the rules supporting its usage.
Is It Necessary for Everyone to Retrieve Data Directly from the Original Server?
Someone already possesses the same data. If they can receive it from that person, there should be no need for everyone to retrieve it directly from the original server.
The image here is similar to traditional P2P file sharing. The person holding the data passes it on to those who need it next.
I want to incorporate a mechanism into this that verifies sources and usage conditions.
Who published the data? Under what conditions can it be redistributed? Is it identical to the original, or if processed, what was it based on? Data should circulate in a state where these can be confirmed.
If done this way, the original provider would not have to bear the burden of distribution alone. By utilizing nearby distribution points or data already acquired, there is room to reduce redundant long-distance transfers.
Of course, simply increasing distribution points does not guarantee a reduction in overall traffic volume. A mechanism encompassing where data is stored and from where it is passed becomes necessary. Nevertheless, I believe there is value in changing the form of "everyone who needs it should come retrieve it from the original site."
The Duties of Users and the Rights of Providers
I am positive about my published information being used by AI itself.
Humans also learn from things created or thought up by others. We do not wish to reject all such usage and imitation.
However, if machines begin collecting, processing, and redistributing in bulk what humans once handled one by one, the scale and how providers perceive it will change. Permitting use is different from saying "any conditions are fine."
What I believe is necessary is making clear the duties of those who use data and the rights of those who provide it.
For users, it becomes necessary to verify under what conditions acquisition, processing, and redistribution are permissible and to use them with peace of mind. If those conditions include source attribution or payment, they must be upheld.
For providers, it is necessary to indicate usage conditions for their data and have it handled according to those conditions. Ideally, I would like them to know how their data is being used and spreading.
What I am considering here is not an explanation of current law where providers can freely restrict all usage, but a proposal for what rules and mechanisms we want in future data circulation.
Collecting without knowing the conditions leaves uncertainty on one side; using without knowing what is happening leaves uncertainty on the other. A mechanism to bridge this gap is needed.
Using Micropayments to Keep Data Moving
Here, micro-payments take on meaning.
For example, paying a very small amount when receiving data. Part of it goes to the original provider, and another part goes to the distributor who actually delivered the data. Could we not consider such a mechanism?
The original provider would have a reason to keep data public. The redistributor would have a reason to use storage capacity and communication lines to deliver it. Users could receive necessary data under approved conditions.
Data that can be used for free should simply circulate under those conditions. There is no need to make everything paid.
What I expect from micro-payments is to increase such options and play a role in supporting circulation.
It is not merely the story of "money goes to the author when read." It is about maintaining an easy-to-use state while sharing the burdens involved in providing and distributing data. I am considering it as a mechanism for that purpose.
Ensuring Traceability Even After Processing
What I want to circulate is not just copies of the original files.
Data is created using certain materials, which are then used for further generation or processing. In such flows, I want records remaining that allow tracing what was based on.
Records of redistributing the same file and records of using it as material for processing/generation must be considered separately. The latter also requires a mechanism where AI and processors record exactly what they actually used.
Recording on a blockchain does not automatically reveal what materials were used. Therefore, beyond the foundation for keeping records, cooperation from those handling data and common rules are also necessary.
Nevertheless, I believe ensuring that sources and usage conditions can be inherited forms the basis for making redistribution easier to approve.
I want to achieve both wide distribution and maintaining the connection with the original provider.
Could This Be Realized with BSV?
As a candidate for this mechanism, what I personally expect is BSV (Bitcoin SV), which is also the blockchain I originally advocated for.
If vast amounts of data pass from person to person and service to service, significant processing capacity will be needed for the records supporting that circulation and micro-payments.
Regarding Teranode, the node software for BSV, the development side reported in May 2024 that over two weeks, they maintained an average of more than one million transactions per second using six nodes distributed worldwide. Official Test Report
This is a result from a test environment and does not mean the data circulation mechanism being considered here has been implemented practically. Nevertheless, if processing at this scale is possible, concepts involving handling vast amounts of circulation records and micro-payments may begin to appear realistic.
The division of roles I envision combines a mechanism for delivering the data itself with a mechanism for verifying its source, usage conditions, and payment. I expect BSV to play the role of supporting these records and payments.
There are points to refine, such as actual costs and methods to verify that data was delivered correctly. BSV is listed as one concrete candidate.
Toward a System That Works Widely and Treats Providers Fairly
If AI begins handling vast amounts of information, forcing everyone to retrieve from the original distribution source each time may become impractical.
Therefore, data should be shareable under approved conditions. The burden of distribution should also be shared. On that basis, verification of the original provider and usage conditions should be possible, with compensation returning as needed.
Users can use it with peace of mind. Providers can state the conditions under which their data may be used. Data will circulate widely while maintaining this relationship.
Distributing the burden on servers and networks while supporting both the duties of users and the rights of providers.
What I expect from blockchain is such usage.
The Conversation Behind This Article
Excerpts and summaries from past conversations related to the points of this article are presented here. These are not verbatim quotes. AI responses are also included as proposals shown within the discussion.
Me: Problems Felt by Those Collecting Data
There is a desire to store data locally as material to pass to AI. Simply collect all the data and decide later how to use it. If everyone starts doing this, I think the load on networks and source servers will become a problem.
Using it as a learning source is fine. However, if data created based on that becomes the source for yet another data, isn't there a need for a mechanism to trace back to the origin?
AI: Proposal to Separate Distribution from Records, and Its Limits
We can consider distributing the data itself separately from records of sources and derivative relationships. If acquisition is possible from a distributor holding the same data, there is no need to retrieve from the original server every time. Blockchain becomes a candidate for supporting records of such lineage.
However, it is necessary to distinguish between original data directly processed and data included in model training. Just because something was used for learning does not make it a direct source citation for individual outputs.
Me: Proposal Linking Secondary Distribution with Micro-payments
I want a mechanism where the author side can know when they are cited. It is conceivable that a very small amount of money enters based on usage. If this is done well, secondary distribution could become a means to reduce the load of data distribution.
AI: Proposal for Distribution and the Option of Free Provision
So you are envisioning delivering data from permitted redistribution points and splitting compensation between the original provider and the distributor? This makes the side performing secondary distribution a participant supporting distribution.
The option to provide for free should also remain. I believe what is central is not the charging itself, but rather that providers can indicate circulation conditions in a form readable by AI as well.
Me: Relationships I Want to Protect with That Mechanism
I do not want to completely prohibit copies. However, I think those being imitated will get confused if they cannot specify conditions. The obligations and rights of the side doing the imitation, and the rights of the side being imitated. Such things seem important to me.
AI: Proposal to Make Usage Conditions Machine-Readable
For users: ability to reference, process, and redistribute under certain conditions. In return: do not erase sources, uphold conditions, and pay compensation if necessary.
For providers: indicate usage conditions and be able to verify how they are used.
This can be organized as a proposal to incorporate such relationships into rules readable by machines.
Revision history
- Initial publication
REFERENCE
Quotations and links are welcome. Full republication is not licensed.
Read the usage policy →
Reader comments
Comments are submitted through Google Forms and appear here only after the author reviews and approves them.
Submit a comment →No published comments yet.
Comment policy