Apologies I’m too scoped in that second paragraph—I assume a specific use case, verified TAIG.
For using an auditor-in-a-box to audit a closed-weight model, it is important that the encrypted closed-weights themselves be submitted* (then subsequently decrypted from within the secure box). This use case was explored here https://www.frameworkzero.org/#architecture, under ‘Vaults’.
*Exception to this might be Attestable.com, however they do not release the relevant open-source to verify their system.
Sure, a zero-trust auditor(s)-in-a-box running an audit on a closed-weight target model, in example:
The Box receives the:
Target Model, encrypted if closed-weight
Capable Models, a.k.a. the auditors
Audit, say a singular model form of Inspect Petri[1]
(Optionally & ideally) AI safety standards that specify limits/red-lines and bans
The Box orchestrates the models in to separate, isolated & attestable Inference Box siblings.
Closed-weights are decrypted from within via [hardware] keys by the owning party.
The Box runs the Audit on the Target Model, coordinating with each Capable ModelInference Box.
The Box releases a privacy-preserving report with attestation.
A TAIG auditor-in-a-box concept works better when appliedbefore training assessment (using a standardize template for architecture, workloads, data, etc) to detect: banned architectures/paradigms, suspicious workloads, dangerous data domains or poisoning, prompt injection, malicious agendas, dangerous engineering, etc.
Both concepts work best as components in a blockchain DAO[2]-based Hybrid-TAIG[3]framework*, serving as verification-in-depth for AI safety standards—and not alignment, or a substitute of.
*Disclaimer: I’m a contributor of the linked website. Not self-promotion here, just concept sharing—a concept that would have more optimally been in a LW post instead of it’s own website.
Hybrid-TAIG: Basically Technical AI Governance that uses slower, consequential human-in-the-loop (as opposed to ‘rubber stamping’ or more ‘human-in-the-loop fatigue’-susceptible ones).
Apologies I’m too scoped in that second paragraph—I assume a specific use case, verified TAIG.
For using an auditor-in-a-box to audit a closed-weight model, it is important that the encrypted closed-weights themselves be submitted* (then subsequently decrypted from within the secure box). This use case was explored here https://www.frameworkzero.org/#architecture, under ‘Vaults’.
*Exception to this might be Attestable.com, however they do not release the relevant open-source to verify their system.
Can you try to dumb it down a bit more for me? Can you try to give a concrete worked example?
Sure, a zero-trust auditor(s)-in-a-box running an audit on a closed-weight target model, in example:
The Box receives the:
Target Model, encrypted if closed-weight
Capable Models, a.k.a. the auditors
Audit, say a singular model form of Inspect Petri[1]
(Optionally & ideally) AI safety standards that specify limits/red-lines and bans
The Box orchestrates the models in to separate, isolated & attestable Inference Box siblings.
Closed-weights are decrypted from within via [hardware] keys by the owning party.
The Box runs the Audit on the Target Model, coordinating with each Capable Model Inference Box.
The Box releases a privacy-preserving report with attestation.
A TAIG auditor-in-a-box concept works better when applied before training assessment (using a standardize template for architecture, workloads, data, etc) to detect: banned architectures/paradigms, suspicious workloads, dangerous data domains or poisoning, prompt injection, malicious agendas, dangerous engineering, etc.
Both concepts work best as components in a blockchain DAO[2]-based Hybrid-TAIG[3] framework*, serving as verification-in-depth for AI safety standards—and not alignment, or a substitute of.
*Disclaimer: I’m a contributor of the linked website. Not self-promotion here, just concept sharing—a concept that would have more optimally been in a LW post instead of it’s own website.
Inspect Petri: an auditing agent that automates monitoring and interaction with LLMs to detect alignment issues and reward hacking. https://github.com/meridianlabs-ai/inspect_petri/
DAO: Decentralized Autonomous Organization (blockchain)
Hybrid-TAIG: Basically Technical AI Governance that uses slower, consequential human-in-the-loop (as opposed to ‘rubber stamping’ or more ‘human-in-the-loop fatigue’-susceptible ones).