Blockchain and the law
What is blockchain?
- An append-only decentralized database
- Maintained by a consensus algorithm
- Stored on multiple nodes (computers)
It has the ability to decentralize business models, forms of human interaction and markets.
It can be used for record keeping, transferring value and smart contracts to automatically execute transactions: #trust.
The legal standpoint and disruptive technologies.
Two normative objectives: fundamental right protection vis-à-vis promotion of innovation.
Blockchain(s)
- On a blockchain, data is usually grouped into blocks that, upon reaching a certain size, are chained to the existing ledger through a hashing process.
- Data is thus chronologically ordered in a manner making it difficult to tamper with information without altering subsequent blocks.
- Hence, immutability of blockchains (at least it is very difficult): then, trust.
- However, it implies that the piece of information was included at some verifiable point in the past, not that the same is correct!
- DLTs (Distributed Ledger Technology) rely on a two-step verification process with asymmetric encryption.
- Every user has a public key (string of letters, username) and a private key (password): the private key can decrypt data that is encrypted through the public key. Public keys hide the identity of the individual unless they are linked to additional identifiers.
Nodes
- The nodes are the computers on which the ledger is stored.
- Some DLTs operate a distinction between "full" and "lightweight" nodes whereby only full nodes store an integral copy of the ledger from the genesis block whereas lightweight nodes only store those parts of the ledger of relevance to them.
- In public and permissionless blockchains, anyone can entertain a node by downloading and running the relevant software. Some (but not all!) nodes also function as "miners", which aggregate transactions into candidate blocks and hash a new block to the chain on the basis of a predetermined consensus protocol. (Bitcoin for example).
- Blockchains can also be private and permissioned, which means that they can run on a private network such as intranet or a VPN (as opposed to the internet) and an administrator needs to grant permission to individuals wanting to maintain a node. (Notarchain for example).
Blockchain(s)
- Public and permissionless blockchain: anyone can participate in the blockchain and become validator.
- Public and permissioned blockchain: everyone can participate in the chain as node and access data, but only some pre-authorized of the participants may act as miners and add data to the ledger.
- Private and permissioned blockchain: both validators and nodes that are merely participants have to be authorized by a set of actors (public and private governance).
GDPR vs Blockchains
- Blockchains offer a record-keeping function that dispenses from the need for third party intermediation (middleman) and by analogy can decentralize the collection, storage and processing of data. This stands in sharp contrast with the current data economy, characterized by economic centralization in the form of 'platform power'.
- Blockchains offer the promise of the decentralized handling of data and data sovereignty, a concept that focuses on giving individuals control over their personal data and allowing them to share such information only with trusted parties.
- GDPR shares the data sovereignty objective as it aims to give natural persons 'control over their own personal data'.
- For examples, the right to data portability (right to receive the data regarding me, and to transfer them to another controller) (Art. 20) enshrines this objective in allowing a data subject to receive data from a controller in order to give it to another controller.
- Blockchains do not, per se, provide any privacy guarantees so that for data sovereignty objectives to be achieved, they must be combined with additional mechanisms. Indeed, despite the technology's promises for data sovereignty, there are also perils for if the necessary safeguards are not implemented; blockchains can reveal any and all data stored on them.
Blockchain does not provide any privacy and guarantee regarding the data, therefore some measures have to be implemented.
- GDPR only applies to 'personal data', defined as 'any information relating to an identified or identifiable natural person' (the 'data subject').
- Where data is rendered completely anonymous, it no longer amounts to personal data and thus falls outside the scope of the legal framework.
- Where data is rendered pseudonymous, however, it continues to qualify as personal data as the indirect identification of a natural personal by an identifier remains possible.
- Two sets of data stored on blockchains can potentially be defined as personal data for the purposes of the GDPR: transactional data stored in the blocks as well as public keys.
Personal Data and Blockchain
- Information stored on blocks may be data related to an identified or identifiable natural person and therefore amount to personal data.
- Data can be stored on a blockchain in three alternative fashions: in plain text, in encrypted form, or by hashing it to the chain.
- Can these processes sufficiently anonymize personal data to allow it to evade the GDPR's scope of application? High threshold for anonymization: the processing must 'irreversibly prevent identification'.
- Personal data stored on a blockchain in plain text clearly remains personal data for the purposes of the GDPR.
- According to the WP29, where data is encrypted it can still be accessed with the correct keys, meaning that it is not irreversibly anonymized (encrypted data can for example be connected to the data subject where transactions are effected for off-chain goods or where crypto-assets are converted into fiat currency).
- Encryption is considered a pseudonymization under the EU data protection regime given that the data subject can still be indirectly identified: so, it is not considered as an anonymization technique.
- Accordingly, transactional data that has been encrypted remains personal data for the purposes of the GDPR.
- The same applies (so far) to hashing.
- Transactional data can therefore be moved off chain to escape from the application of the GDPR
- Public keys? They are pseudonymous data (thus subject to the GDPR) and cannot be moved off chain as quintessential component of the technology.
(GDPR is still applicable on pseudo-anonymized data because even if the data is encrypted through hash, it isn't an irreversible process, there is still the possibility to go back to decrypted data). The only way to escape GDPR is to move data off chain.
Data Controller
- When it comes to private blockchains, it might still be possible to identify a central intermediary that can qualify as the data controller such as the systems operator that will be the addressee of the data subject's claims.
- For other DLTs, there is no central point of control as the network is operated by all nodes in a decentralized fashion.
- Permissionless blockchains are distributed and decentralized peer-to-peer networks that everyone can participate in to interact with unknown or untrusted counterparties. In such a setting either no node qualifies as the data controller in the absence of independent determination of the means and purposes of processing, or, more likely, every node qualifies as a data controller.
- Nodes are indeed not subject to external instructions, autonomously decide whether to join the chain, and pursue their own objectives.
- Determining that each node is a data controller raises considerable complications. The exact number, location and identity of nodes on a chain cannot be established without difficulty.
- Nodes are furthermore passive entities subject to the directions of software designed by developers. They (i) only see the encrypted or hashed version of the data; and (ii) are unable to make any changes thereto. Nodes are thus decentralized entities that cannot respond to the tasks the GDPR requires of centralized agents.
- The enforcement of obligations resting on nodes is thus burdened by significant difficulty.
- For the Bitcoin blockchain, there are currently approximately 11,000 nodes around the planet, of which about 1800 are in Germany and 800 in France.
- If one were to address each of these nodes, some of which may not be found in a single jurisdiction this would create two sets of problems.
- First, a large amount of nodes would need to be contacted and compelled to comply, as opposed to a single controller in a data silo scenario.
- Second, this may lead to forcing all nodes to stop running the blockchain software where GDPR rights cannot be achieved through alternative means.
- This would result in a situation where an entire blockchain would be taken down in one jurisdiction for noncompliance with a single data subject's rights, which may be considered disproportionate.
- It is moreover unclear how fines will be calculated where a data controller on an unpermissioned blockchain has failed to comply with data protection requirements given that Article 83 GDPR calculates them on the basis of annual worldwide turnover.
- French CNIL (Commission Nationale Informatique & Libertès): participants who have the right to write on the chain and who decide to send data for validation by the miners can be considered as data controllers. In fact, blockchain participants define the purposes and means of the processing. (e.g., if a notary records his/her client's property deed on a blockchain, he/she is a data controller. If a bank enters its clients' data onto a blockchian as part of its client management processing, it is a data controller).
Territorial Scope of Application
- Unpermissioned blockchains usually run on nodes located in various jurisdictions across the globe, leaving creators with no control over the geographic spread of the network. This makes DLTs inherently transnational in nature, triggering a range of jurisdictional issues.
- GDPR applies 'to the processing of personal data in the context of the activities of an establishment of a controller or processor in the European Union, regardless of whether the processing takes place in the Union or not'.
- The GDPR's broad territorial scope accordingly likely entails that its obligations bind many blockchain-based applications with only an indirect link to the EU.
- A further jurisdictional question relates to the application of European data protection requirements to the transfer of data to third countries.
- On permissionless ledgers we can presume that there is always an element of cross-border data processing: the data stored in blocks is hashed to the chain by a randomly selected miner that can be based anywhere. The ledger is subsequently updated on each node to reflect the addition of the new block.
- Transfer to third countries is possible under certain conditions (possibility of a data subject providing explicit consent for such a transfer, subject to being informed about possible risks). This could be easily implemented on a private blockchain where access is controlled and can be subjected to terms and conditions but it is not obvious how such consent could be acquired in respect of a permissionless chain.
Enforcement of the GDPR rights
- While from a legal perspective a data subject can invoke her rights vis-à-vis every single node, it is far from obvious how, from a technical perspective, nodes could implement related requests to correct, erase or restrict data.
- How a data subject can consent to the processing of his/her personal data on a blockchain indeed remains an as of yet unresolved question.
- Data Minimization
- The GDPR mandates that personal data be 'collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes'.
- Conversely, once added to a blockchain, data will perpetually remain part of the chain, given that it is an append-only database that continuously expands. Distributed ledgers are by definition ever-growing creatures, which augment and accumulate further data with each additional block.
- Integral copies of the chain are stored on each full node, quite the opposite of the data minimization spirit. Once data has been added to the chain, it can in principle no longer be amended or deleted, which makes it difficult not to say impossible to implement the minimization principle and storage limitation requirements.
- Right to amendment/rectification
- The GDPR requires that personal data be accurate and up to date. Where this is not the case, 'every reasonable step must be taken to ensure that personal data that are inaccurate, having regard to the purposes for which they are processed, are erased or rectified without delay'.
- Two practical impasses:
- A data subject cannot possibly identify any or all of a blockchain's full nodes.
- Second, even if the data subject succeeds in addressing a claim under the GDPR, nodes are simply unable to change any of the encrypted data stored in a block (immutability of ledgers).
- Personal data can also be rectified 'by means of providing a supplementary statement'.
- Could the addition of new data to the chain of blocks, which rectifies data previously added (without however deleting the original entry), be considered to comply with the requirements set by the GDPR?
- The GDPR also requires that the controller communicate any rectification or erasure of personal data to 'each recipient to whom the personal data have been disclosed'. This, can however be presumed to not apply to nodes as the same provision clarifies that controllers are dispensed from said obligation where 'this provides impossible or involves disproportionate effort'.
- Right to access
- A data subject has the right to obtain confirmation from the controller whether or not her personal data is being processed. Data subjects are moreover entitled to be informed about safeguards that apply where data is transferred to third countries (relevant for blockchain, as a node validating a block in the EU will thereafter share that information with all nodes of the blockchain, irrespective of their geographical location).
- Controllers don't know which data is stored on the blockchain as they often only handle the encrypted or hashed version thereof. Even if a data subject were successful in contacting a node, the latter would be incapable of verifying whether a data subject's personal data is being processed. The data subject could of course join an unpermissioned network and obtain a copy of all data, but it is questionable whether this would be regarded as a satisfactory solution in the eyes of the GDPR.
- Article 15(3) GDPR moreover entitles data subjects to obtain a copy of their personal data undergoing processing from controllers, which would be equally impossible where its has been cryptographically pseudonymized.
- Storing personal data off-chain is to be preferred for transactional data but remains unfeasible for public keys.
- Right to be forgotten
- Immutability is one of blockchains' most heralded features. They are, by definition, unable to forget as they were specifically designed to be censorship-resistant. A straightforward application of the right to be forgotten to DLTs can be excluded.
- However, the notion of 'deletion' may refer to a variety of options so that requests of erasure may be fullfilled in different manners.
- It is worth distinguishing between transactional data and public keys.
- Different solutions may apply to transactional data. Where personal data is recorded in a referenced encrypted and modifiable database as opposed to the blockchain itself, it can be deleted in line with data protection requirements without the need to touch the blockchain.
- With regard to public keys compliance is again more burdensome. It must be recalled that the right to be forgotten is not an absolute right. Article 17(2) GDPR rather provides that the controller shall take 'account of available technology and the cost of implementation'. Could the reference to 'available technology' lead to an interpretation of the GDPR that dispenses from outright erasure in light of blockchains' technical limitations in favour of an alternative solution?
Blockchain and the GDPR
- Whereas the GDPR was fashioned for an age of centralized data silos, blockchains promise a future of decentralized data management.
- The new supranational data protection framework is already partly outdated in respect of its application to distributed ledgers for it simply cannot account for the technology's characterizing features.
- The same conclusion has been reached in respect of big data and artificial intelligence, indicating considerable challenges ahead.
- However, blockchains, if adequately designed, and the GDPR can share a common objective: giving a data subject more control over his/her data. This is of course only the case where blockchains are specifically fashioned to achieve that objective.
Artificial Intelligence
There are many issues related to AI and which are similar to the one relative to blockchain. Those are not only related to the GDPR even though the processing of personal data in artificial intelligence systems is a very debated issue. There are many opinions of this since the automated decision-making processing is something which is developing rapidly, and for this reason in Europe there is a specific provision preventing entirely automated decision systems to affect human rights. (art. 32 of the GDPR). Therefore, algorithms and AI might entail the processing of data and create some challenges.
However, the most debate issue regard civil liabilities connected to the damages caused by the usage of AI systems. But there is not a provision giving specific rules.
Even if more and more machines will learn how to carry out different activities risks are always involved not only on the instruction received but also regarding the inputs that the machines receive.
The challenge is to make sure that AI systems are working more like men rather than "animals"/resilient machines, they should develop skills and new functions.
The problems arise when the source of the input of the machine is not known or if the input is based on the experience of the AI machine.
Then, who is responsible of the harm in front of the law for an AI system?
- The manufacturer (the programmer): he is including some inputs but the harm might occurs not because of those inputs. Maybe from the inputs received during the activity
- The user: difficult to accuse a user
- The insurance company
It depends also which tradeoff between human right protection and innovation we want to follow:
- protection of human → strict rules and discouraging innovation
- soft rules → no incentives for the manufacturers to comply with very high standard
In general, there are two possible options:
- Attaching "electronic personality" to AI systems requiring the owner entering into insurance contract (like in the Ancient Rome for slaves, to whom a peculium was attached). Therefore the owner (which might not be the user) is liable. And he normally buys an insurance to cover the possible danger carried out by the AI system.
- Regulating "ex novo" liability of AI systems: creating new rules. The possible new rules are product liability.
Product liability might be based on
- Negligence: focusing on the conduct of the manufacturer; liability occurs in case of violation of a duty of care which actually caused a harm. You are seeking for an element in the conduct of the manufacturer which justices the fact that the cost of the harm is charged to him.
- Strict liability: strict liability claims focus on the product itself rather than on the conduct of the manufacturer; the manufacturer is liable if the product is defective, even if the manufacturer was not negligent in making that product defective.
Both options can be considered controversial. With the negligence option you are discouraging those who wants to use and implement AI systems since you are liable only if you violate a duty of care and you undermine the protection of the possible victims. In the other case you are providing the strongest protection to victims, so the manufacturer has no incentives to implement a AI system since he is always liable even if it is not his fault.
another problem: Machine Learning
It is not easy to allocate liabilities when you do not have something which is not executed as a result of an input given the manufacture. There is a divergence between the preconstructed behaviour and the behaviour experience, it can create harm.
The users should have the responsibility of making sure that the machine is learning in a proper way.
But there are also some liability regimes proposed by experts/commentators :
- Strict product liability (Directive 374/85/CEE)
- Tort liability (based on negligence/intention)
- Parental liability (objective liability since you are supposed to look after the action of the system)
- Liability for dangerous activities (burden of proof is reversed), or liability for animals
- Legal subjectivity (criminal law)
Those are all possible options, all with some pros and cons.