The tools I use, and what their makers publish
Enquiry run: 1 September 2026. All sources checked on that date. Next review: September 2027.
This is the record required by commitments 1, 2, 4 and 5 of my Standards of Practice, and by test 5. It names every generative tool I use, sets out what each vendor actually publishes about how it was built, who labelled its data and what it costs environmentally — and names, per field, what I could not establish.
How to read it. Where a cell says docs silent, I went to the named document, read what was in it, and it was not there — the document and the date are given so anyone can check. Where a cell says not established, either no document exists or I could not retrieve one, and the cell says which. I have not written "unclear" anywhere, because that hides whether anyone looked.
Nothing in here is an estimate. Where I could not find something, it says so.
The honest paragraph
Almost nothing about how these tools were made is published by the people who made them.
Not one of the seven vendors I use publishes what the people who labelled its training data are paid, where they work, who employs them, or on what terms. Two of them — Google and Anthropic — confirm in passing that such work happens; neither says who does it. That is the whole of the industry's disclosure on the human cost of these tools, and it is why my fourth commitment names a vacancy rather than promising an action. There is nothing here to choose between. No creative AI product I use competes on the treatment of the people who annotated its data, because none of them will say.
On what the models were trained on, the position is barely better. Every one of the seven describes its training data in categories — "publicly available information", "public and private datasets" — and not one names a source, a dataset, a date range, or how much of the corpus came from where. Two of them contradict themselves: Anthropic's own two official descriptions of what its flagship model was trained on do not match, and Google's video model card describes at length how its footage was filtered and captioned while never once saying where the footage came from. Google has confirmed publicly, through a spokesperson, that it trains on YouTube videos; that confirmation appears in journalism and in none of its own documents. There is no tool in my stack that is opt-in trained, and none with a licensed corpus. I did not choose scraped tools over licensed ones — for image and video generation at this scale, in September 2026, no licensed or opt-in alternative existed to choose.
On the environmental side there is a real split, and it is not the one you would expect. It has nothing to do with how big or how ethical a company is. It is simply whether they own the buildings and whether somebody makes them report. Google owns its datacentres and publishes an annual environmental report: a fleet power efficiency figure of 1.09, per-campus figures for around 35 named sites, and — creditably — its own hourly clean-energy data showing that the same company running at 100% carbon-free power in Stockholm runs at 1% in Hong Kong. Kuaishou, which makes Kling, is listed in Hong Kong and files a mandatory ESG report: a power efficiency figure of 1.20 and 93% green electricity for its self-built datacentres. Those are real numbers, honestly published, and both companies deserve credit for publishing them.
OpenAI, Anthropic and Runway publish nothing at all. No sustainability report, no environmental page — in each case a literal missing page where one would sit. Runway's compute supplier, CoreWeave, has not yet reported even its own basic carbon emissions and says the efficiency data belongs to the landlords it rents from, whom it does not name. Anthropic went furthest: asked by Stanford researchers for its energy, carbon and water figures, it answered that the information is "proprietary and not disclosed publicly to protect competitive advantages and intellectual property." That is not an oversight. It is a decision, on the record.
And then there is the thing that undoes even the good disclosure. Not one of these seven tells you which datacentre served your request. Google's terms say your data may be cached "in any country in which Google or its agents maintain facilities". OpenAI offers a region guarantee only to enterprise customers, in three regions, and says even that "does not confine all processing to the region". So Google's excellent per-campus numbers, and Kuaishou's real ones, cannot honestly be attached to anything I actually made. I know the fleet average. I do not know, and cannot find out, which building rendered my clip. Anyone who quotes a company-wide figure as the footprint of a particular piece of AI work is guessing, and I am not going to do that.
The one figure that exists for a single prompt — Google's 0.24 watt-hours — is for a text prompt, and it explicitly excludes the energy of training the model in the first place. There is no published figure anywhere in this stack for the cost of generating one image or one second of video. Those are the two things I actually do.
Four of the seven use what I type and upload to train their models by default, and I have to go and turn that off. On Kling, the only way to withdraw is to send an email, and the wording covers future use — it says nothing about material already trained. On Runway there is no way to turn it off at all: the terms claim the right to train on inputs and outputs, and the document contains no opt-out clause anywhere, though it does contain a thirty-day opt-out of arbitration. What Runway's terms do say, and it is worth saying plainly because the adjectives are alarming out of context, is that the licence is limited — to training, moderation, and operating the service. Runway does not claim to own my inputs or my outputs and does not restrict my commercial use of them. It is a licence to learn from what I make, not a licence to use it. And paying does not fix it: a paid consumer subscription to Google or OpenAI buys exactly the same training terms as the free tier. Only business and API accounts are excluded.
I am still using all seven. Until I am generating in house, that is the tent I live in and the variables I have to accept. I am telling you what I could not find out, because that is the whole of what the duty of enquiry claims — an act, not an outcome.
Summary — the seven tools at a glance
Wide table — scroll it sideways →
| Tool | What I use it for | Training data | Data labour | Environmental | Trains on my content? |
|---|---|---|---|---|---|
| Nano Banana (Google) | stills, first frames | scraped — crawling disclosed, no source named |
docs silent | fleet figures only, unattributable | via Gemini terms |
| Veo / Flow (Google) | short clips | undisclosed — no source named in any Google document |
docs silent | fleet figures only; no video figure exists | via Gemini terms |
| Kling (Kuaishou) | main video generation | scraped ⚠ single source |
docs silent | PUE 1.20, 93% green — for buildings in China, not for this | Yes — opt-out by email only |
| Runway (Runway AI) | video | undisclosed |
docs silent — complete | nothing published, at any link | Yes — no opt-out exists; licence purpose-scoped to training and operating the service |
| Gemini (Google) | writing | scraped |
docs silent | text prompt: 0.24 Wh, excl. training | Yes — on by default |
| ChatGPT (OpenAI) | writing | undisclosed |
docs silent (vendor) | nothing published | Yes — on by default |
| Claude (Anthropic) | writing | scraped |
docs silent | declined as proprietary | Yes — opt-out per terms |
Consent basis, allowed values: opt-in trained · licensed corpus · scraped · undisclosed. No tool in my stack is opt-in trained or licensed corpus.
The full record, tool by tool
Every claim below carries its source and the date I checked it.
1 · Nano Banana — Google
| Field | Record |
|---|---|
| Tool, vendor, model/version | Nano Banana, Google. Current IDs: gemini-3.1-flash-image (Nano Banana 2), gemini-3-pro-image (Pro), gemini-3.1-flash-lite-image (2 Lite). ⚠ Unqualified "Nano Banana" now denotes gemini-2.5-flash-image, which Google's own docs mark for transition away. ai.google.dev/gemini-api/docs/image-generation, 1 Sep 2026 |
| What I use it for | Stills and first frames |
| Training data — what IS published | The image model cards contain no training data content; they defer twice, ending at the Gemini 3 Pro model card, which states the corpus includes "publicly available datasets that are readily downloadable; data obtained by crawlers; licensed data obtained via commercial licensing agreements; user data …; other datasets that Google acquires or generates in the course of its business operations, or directly from its workforce; and AI-generated synthetic data." Gemini 3 Pro Model Card, Nov 2025 upd. May 2026, 1 Sep 2026 |
| Consent basis | scraped — crawling affirmatively disclosed; no opt-in mechanism described for crawled content |
| Data labour | docs silent — enumerated all seven sections of four model cards (3.1 Flash Image, 3 Flash, 3 Pro, 3 Pro Image). "Human raters", "annotators" and "crowdworkers" do not appear in the Gemini 3 Pro card. It discloses "human-preference data" as a training input and says nothing about who produced it. 1 Sep 2026 |
| Environmental — WUE | Not established. Google's 2026 Environmental Report is 112 pages; retrieval failed past ~p.38 across five attempts on two hosts, and the data appendix at ~p.82 was never reached. WUE appears nowhere in the retrievable narrative, nowhere in the announcement blog, and nowhere on Google's data-centre pages — and Google publishes per-site PUE while publishing no per-site WUE. I cannot state that Google publishes no WUE. I can state I could not reach the appendix. |
| Environmental — PUE | 1.09, global fleet, 2025. "In 2025, the average annual power usage effectiveness for our global fleet of data centers was 1.09." Per-campus figures published for ~35+ named sites (Central Ohio 1.04 · The Dalles 1.06 · Dublin 1.08 · Singapore 1.14). datacenters.google/efficiency, 1 Sep 2026 |
| Environmental — water source | docs silent. Google publishes total consumption — "In 2025, we consumed 10.9 billion gallons (41 billion liters) of water across our data centers and offices" — and water-stress risk ("87% of our freshwater withdrawal came from sources with low or medium risk"). It does not disclose water source type at any level of granularity. 2026 Environmental Report; datacenters.google/operating-sustainably, 1 Sep 2026 |
| Environmental — heat reuse | Published, two named sites. Finland: "recovered heat from Google's data center is expected to cover 80% of the annual heat needs" (⚠ prospective, not measured). Germany: district heating collaboration. No fleet-wide figure. 2026 Environmental Report p.26, 1 Sep 2026 |
| Environmental — renewables | ANNUALLY matched, stated as such by Google. "we again matched 100% of our electricity consumption with renewable energy purchases (on a global and annual basis) for the ninth consecutive year." Google separately defines the harder standard it has not met: "24/7 CFE requires an hourly and local match." Its own hourly Cloud data ranges from 1% (Hong Kong) to 100% (Stockholm). 2026 Environmental Report pp.5, 27; cloud.google.com/sustainability/region-carbon, 1 Sep 2026 |
| 🔴 What could not be established | Any image-specific training provenance — the chain ends in a text-model card. Corpus size, crawl date range, named licensors, proportional split between the six disclosed categories. Google's copyright or opt-out position — the words "copyright", "opt-out" and "consent" do not appear in the Gemini 3 Pro model card. Any energy, water or carbon figure for image generation — the section titled "Implementation and Sustainability" in all four image model cards contains hardware prose and no metric. Fleet WUE (see above). Water source type. Which region serves a request. |
| Sources | ai.google.dev/gemini-api/docs/image-generation · Gemini 3 Pro / 3 Flash / 3.1 Flash Image / 3 Pro Image model cards · sustainability.google 2026 Environmental Report · datacenters.google/efficiency · datacenters.google/operating-sustainably · cloud.google.com/sustainability/region-carbon — all 1 Sep 2026 |
2 · Veo and Flow — Google
| Field | Record |
|---|---|
| Tool, vendor, model/version | Veo 3.1, Google DeepMind, used via Flow (labs.google/flow). ⚠ The Gemini API video docs name "Veo 3.1" but print no model ID, unlike the image docs. 1 Sep 2026 |
| What I use it for | Short social-tier clips |
| Training data — what IS published | The complete disclosure, from the terminal document: "Veo 3 was trained on audio, video, and image data. Audio and video datasets were annotated with text captions at different levels of detail, leveraging multiple Gemini models, and filtered to remove unsafe captions and personally identifiable information." The Veo 3 Tech Report adds only: "We train on a large dataset comprising images, videos, and associated annotations." The words "YouTube", "licensed", "publicly available" and "copyright" appear in neither document. Veo 3 Model Card; Veo 3 Tech Report, 1 Sep 2026 |
| Consent basis | undisclosed — no source is named in any Google document. ⚠ But Google has confirmed the source publicly, outside its documentation. CNBC (19 Jun 2025) reported Google confirmed it uses YouTube videos to train Veo 3; a YouTube spokesperson: "We've always used YouTube content to make our products better, and this hasn't changed with the advent of AI." Creators can opt out for select third-party AI companies — not for Google. This is journalism, not vendor disclosure, and both facts belong on the record. |
| Data labour | docs silent — enumerated the Veo 3 and Veo 3.1 Lite model cards (7 sections each) and all 23 sections of the Veo 3 Tech Report. Captioning is disclosed as machine-generated by Gemini models. The human labour behind safety filtering and compliance review is not described. 1 Sep 2026 |
| Environmental — WUE / PUE / water source / heat reuse / renewables | As Google, above — the same fleet figures, the same limits. The Veo 3 Tech Report contains no compute, energy or environmental information across all 23 sections. |
| 🔴 What could not be established | Where the training video came from — the two terminal documents describe how it was filtered and captioned and never say its origin. The YouTube licence-grant language (youtube.com/t/terms is robots.txt-disallowed to my tools). The rights-clearance basis for the video corpus. Any energy or water figure for video generation — none exists. Which image model version Flow runs (it says "Nano Banana" without a version). "Gemini Omni", Flow's top-listed video model: no Google model card, technical report or documentation page could be located at all. Flow's own position on training data and uploads — labs.google/flow/about addresses none of it. Serving region. |
| Sources | Veo 3 Model Card · Veo 3 Tech Report · Veo 3.1 Lite Model Card · deepmind.google/models/veo · labs.google/flow/about · ai.google.dev/gemini-api/docs/video · CNBC 19 Jun 2025 (via syndication — re-source before quoting) — all 1 Sep 2026 |
3 · Kling — Kuaishou Technology
| Field | Record |
|---|---|
| Tool, vendor, model/version | Kling AI, Kling 3.0 family (Video 3.0, Video 3.0 Omni, Image 3.0, Image 3.0 Omni), announced 5 Feb 2026. ⚠ The service is operated by Kling AI Pte. Ltd. (Singapore). The ESG report is filed by Kuaishou Technology (HKEX 01024). They are not the same entity, and the Singapore entity publishes no environmental disclosure of any kind. 1 Sep 2026 |
| What I use it for | Main video generation — keeper and showcase pieces |
| Training data — what IS published | Nothing, in any document I contract under. Enumerated the Terms of Service and Privacy Policy in full; neither contains any statement about what the models were trained on. The only vendor description of data acquisition is in a research paper: "an automated pipeline for large-scale internet data mining". ⚠ That paper is dated Dec 2025 and describes Kling-Omni; Kling 3.0 shipped Feb 2026. kling.ai/docs/user-policy; arXiv 2512.16776, 1 Sep 2026 |
| Consent basis | scraped ⚠ SINGLE SOURCE — one sentence, in one research paper, covering one model, published two months before the version I use. A stricter reading would say undisclosed for shipped Kling 3.0. |
| 🔴 Does it train on my content? | Yes, and it is disclosed. ToS §4.7.3(f) permits use of my content to "create, test, improve, train, or otherwise develop the artificial intelligence or machine learning models, systems, architecture, weights or related technology used by Kling AI". "Content" is defined to include both my inputs and my outputs. The only opt-out is §4.7.4: "you may notify us to revoke the authorization by sending an email to support@kling.ai." No in-product control. The wording is "continue using" — it addresses future use and says nothing about material already trained. ⭐ Note the clause is already purpose-limited on its face — the permitted use is to "create, test, improve, train, or otherwise develop" the models "used by Kling AI". It is a licence to develop models, not to exploit the content. ⚠️ NOT CHECKED, carried forward: whether Kling's terms also contain a separate general content licence alongside §4.7.3(f), of the kind Runway carries at §4.3. Runway's two clauses do different jobs and only one is about training; the equivalent enumeration has not been run on Kling, so no claim is made either way. 1 Sep 2026 |
| Data labour | docs silent — searched the full 2025 ESG Report for annotation, labelling, annotator, 标注, content moderation, outsourced, contractor, labour rights. No passage addresses the conditions of people who annotate AI training data. ⚠ The report does state "Over 6,000 people participated in content review training" — those are platform content reviewers, not AI training-data annotators, and no employment status, conditions or pay is given. I am not counting it as a data-labour disclosure. 1 Sep 2026 |
| Environmental — WUE | docs silent — and this is the sharpest finding here. The 2025 ESG Report refers to "improving the WUE and PUE of data centers" (p.36) and publishes no WUE value for any facility, any year. The vendor names the metric and does not give it. |
| Environmental — PUE | 1.20 average, 1.14 minimum, 2025. "The average PUE value of Kuaishou's self-built data centers in 2025 was 1.20, with a minimum achievable value of 1.14." ⚠ Self-built only — the report elsewhere refers to "self-built and leased" datacentres and gives no figure for leased capacity. Kuaishou 2025 ESG Report, pub. 24 Apr 2026, checked 1 Sep 2026 |
| Environmental — water source | Offices only, and only partially. Beijing HQ "has completed the introduction of municipal reclaimed water, which is expected to cover approximately 20% of the total water demand of non-toiletry water in toilets." For datacentres, an intention only. No data-centre water withdrawal or source is disclosed. |
| Environmental — heat reuse | docs silent — enumerated the environmental chapter and searched the whole document for waste heat, heat recovery, heat reuse and district heating. No scheme is mentioned. "Heat" appears only as "purchased heat", an energy input. |
| Environmental — renewables | 93.0%, annual basis, certificate-backed — and no matching basis is stated anywhere. "In 2025, Kuaishou's self-built data center purchased a total of 583,720.0 MWh of green electricity, accounting for 93.0% of its total electricity consumption", via "green power trading and green power certificate procurement". Green power certificates are annual accounting. The report never claims hourly matching. That absence is the finding — the 93% does not establish that my clips were rendered on renewable power at any given hour. |
| 🔴 What could not be established | The training corpus for the version I actually use. Dataset names, sizes, licensing, provenance. Whether any third-party rightsholder can opt out — §4.7.4 gives revocation to users only. Whether the opt-out affects already-trained weights. WUE (named, unquantified). Data-centre water withdrawal and source. Heat reuse. PUE for leased datacentres. ⚠ The 2025 report's Key Environmental Indicators appendix — four failed fetches, download proxy-blocked. Which cloud provider hosts kling.ai's Singapore servers — the Privacy Policy names none, so the chain terminates with nothing to attribute to. Where Kling inference physically runs. Any environmental disclosure by the operating entity. |
| Sources | kling.ai/docs/user-policy · kling.ai/docs/privacy-policy · Kuaishou Technology 2025 ESG Report (ir.kuaishou.com) · Kuaishou 2024 ESG Report · arXiv 2512.16776 · ir.kuaishou.com Kling 3.0 launch release — all 1 Sep 2026 |
4 · Runway — Runway AI, Inc.
| Field | Record |
|---|---|
| Tool, vendor, model/version | Runway, Gen-4.5 (also Aleph 2.0, GWM-1, Act-Two). ⚠ Runway also serves third-party models inside its own product — Seedance 2.5 (ByteDance) and GPT Image 2 (OpenAI) — which carry their own separate, undisclosed provenance. An output made in Runway is not necessarily a Runway model output. 1 Sep 2026 |
| What I use it for | Video generation. Live paid subscription; generation temporarily paused pending an account dispute |
| Training data — what IS published | One sentence, from 2024, about a superseded model. Runway's CTO to TechCrunch (17 Jun 2024): "We have an in-house research team that oversees all of our training and we use curated, internal datasets to train our models." TechCrunch adds Runway "wouldn't say" where the data came from. Enumerated and found silent on corpus: Terms of Use (18 sections), Privacy Policy (15), Safety page, About page, Gen-4.5 research page, API docs (9 sections). The Gen-4.5 page confirms pre-training happened and names no dataset. The Safety page says data is "filter[ed] before training" and never says what is filtered. 1 Sep 2026 |
| Consent basis | undisclosed. Not licensed — the one licensing deal on record (Lionsgate, 2024) produced a bespoke customised model and says nothing about Gen-4.5. Not scraped — that is alleged in a pending complaint (Businessing LLC v. Runway AI, S.D.N.Y., filed 27 Feb 2026) and reported by 404 Media, and an allegation is not a finding of fact. I am not recording an unproven allegation as an established consent basis. |
| 🔴 Does it train on my content? | Yes, and no opt-out exists — but both licences are purpose-scoped, and the scope is the point. §4.4 first states: "The Company does not claim ownership of any of your Inputs or Outputs. Subject to your compliance with the Agreement, the Company does not restrict your commercial use of your Outputs." It then states Inputs and Outputs "may be used by the Company to train and improve its AI models, algorithms and related technology, products and services (including for labeling, classification, content moderation and model training purposes)", and grants "a non-exclusive, irrevocable, perpetual, worldwide, royalty-free, fully paid, transferable, sublicensable right and license to use any Inputs and Outputs Made Available by you or otherwise generated in connection with your use of the Services at any point, in connection with the purposes described above." ⚠️ CORRECTION, 1 Sep 2026: an earlier draft of this table attributed that licence to §4.3 and quoted its adjectives without its purpose clause. It is §4.4's licence, and it is limited to the purposes above. It is not a general exploitation licence. §4.3 is a separate and ordinary platform licence: "Subject to any applicable account settings that you select, you grant Company a fully paid, royalty-free, perpetual, irrevocable, worldwide, non-exclusive and fully sublicensable right … to host, use, license, distribute, reproduce, modify, adapt, publicly perform, and publicly display … Your Content (in whole or in part) for the purposes of operating and providing the Services to you and to our other users." ⭐ The precise finding is an asymmetry, not an absence: §4.3 is expressly "subject to any applicable account settings that you select"; §4.4 contains no equivalent clause. The content licence bends to settings; the training licence, as written, does not. §16.9 provides a thirty-day opt-out of arbitration. Verified verbatim against runway.com/terms-of-use, 1 Sep 2026 |
| Data labour | docs silent — and the silence is complete. Enumerated Terms (18 sections), Privacy Policy (15), Safety (3), About, careers listings, help centre. §4.4 names "labeling, classification, content moderation" as purposes and identifies no workers, no vendor, no conditions. Searched specifically for Scale AI, Surge, Appen and Sama: no evidence links Runway to any named data-labelling vendor. There is also no named investigative journalism or peer-reviewed work about data labour at Runway. 1 Sep 2026 |
| Environmental — all five metrics | Nothing published, at any link in the chain. runway.com/sustainability returns HTTP 404. Enumerated for any energy, emissions, carbon, water or climate content: Terms (18), Privacy Policy (15), Safety (3), About, Gen-4.5 page, help-centre security article, API docs (9). None contains any. No B-Corp status, climate pledge or SBTi commitment. |
| The chain, and where it breaks | Runway → CoreWeave → unnamed landlord operators → unknown region. CoreWeave's own release (11 Dec 2025) confirms it powers Runway's models "for training and inference". CoreWeave publishes no PUE figure — "We are collecting the necessary PUE data", and it comes from the operators it rents from. No WUE — the term does not appear in its FY2025 10-K. No water source. Heat reuse asserted with no site, counterparty or quantity. No renewable percentage on any stated basis — its strongest claim is that "several" leased datacentres are "powered by non-emitting sources of energy", a term it does not define. 🔴 CoreWeave has not yet reported its own Scope 1 or Scope 2 emissions: "We are currently evaluating and measuring Scopes 1 and 2 GHG emissions". And CoreWeave names no landlord and no site except one joint venture in Kenilworth, New Jersey. CoreWeave FY2025 10-K; coreweave.com, 1 Sep 2026 |
| 🔴 What could not be established | The training corpus of any Runway model. Whether a paid contract overrides §4.4 — no enterprise agreement is public. Any training opt-out at all. Runway's subprocessor list (/subprocessors returns 404). Serving region. The workload split between CoreWeave and AWS. All five environmental metrics, at all three links. Any Runway account of its own training data in litigation. ⚠ Four pages under runway.com describing a different company's financial product were excluded from evidence; a "Runway runs on Google Cloud" claim sourced from runway.com/security should not be repeated. |
| Sources | runway.com/terms-of-use · runway.com/privacy-policy · runway.com/research/introducing-runway-gen-4.5 · runway.com/safety · docs.dev.runwayml.com · help.runwayml.com Enterprise FAQ · runway.com/sustainability (404) · coreweave.com/news CoreWeave–Runway agreement · CoreWeave FY2025 Form 10-K (SEC) · TechCrunch 17 Jun 2024 · Businessing LLC v. Runway AI complaint — all 1 Sep 2026 |
5 · Gemini — Google (as a writing tool)
| Field | Record |
|---|---|
| Tool, vendor, model/version | Gemini, Google. ⚠ Which model wrote a given piece cannot be established from vendor documentation. The app publishes a model picker — 3.7 Flash · 3.5 Flash-Lite · 3.1 Pro · 3.1 Deep Think — and never names a default. The release-notes page's last text-model entry is 3.6 Flash (21 Jul 2026), a model replaced in the picker in mid-August per journalism. The vendor's own changelog is behind its own product. 1 Sep 2026 |
| What I use it for | Writing — scripts and prose |
| Training data — what IS published | Inherited from the Gemini 3 Pro model card (see Nano Banana above). No model card exists for 3.7 Flash, 3.5 Flash-Lite or 3.1 Deep Think. The Gemini 3.1 Pro card contains no training-data section at all; it defers to the Nov 2025 card. |
| Consent basis | scraped — for the foundation corpus. For my own chats on the consumer app, free or paid: also scraped, in the sense that the affirmative act available to me is refusal, not permission. |
| 🔴 Does it train on my content? | Yes, on by default. "If you're 18 or over, Keep Activity is on by default." / "Google uses this data to provide, develop, and improve its services (including training generative AI models)." Turning it off works — with two carve-outs: "won't be used to train our AI models, unless you choose to send Google feedback", and "Chats reviewed by human reviewers … are not deleted when you delete your activity. Instead, they are retained for up to three years." Uploads, images, videos and shared screens are treated as prompts. A paid Google AI Pro/Ultra subscription changes nothing — only Workspace and the paid API are excluded. 1 Sep 2026 |
| Data labour | docs silent, with one line of acknowledgement. "A subset of chats are reviewed by human reviewers (including Google's trained service providers) to help improve Google services." / "Reviewed chats are retained for up to three years." Who they are, where they work and what they are paid: nothing. Google's position on the workers, quoted from reporting on the GlobalLogic rater layoffs: "As the employers, GlobalLogic and their subcontractors are responsible for the employment and working conditions of their employees." Gemini Apps Privacy Notice; WIRED via syndication, 1 Sep 2026 |
| Environmental — the one real per-prompt figure, and its limits | Text prompts only. Google: "the median Gemini Apps text prompt consumes 0.24 Wh of energy … and consumes the equivalent of five drops of water (0.26 mL)"; and separately, "a median Gemini Apps text prompt generates 0.03 gCO2e and consumes 0.26 mL of water". It explicitly excludes training — "LLM training & data storage: This study specifically considers the inference and serving energy consumption of an AI prompt" — plus external networking and end-user devices. Data is from May 2025. Google, "Measuring the environmental impact of delivering AI at Google Scale", 21 Aug 2025, checked 1 Sep 2026 |
| Environmental — WUE / PUE / water source / heat reuse / renewables | As Google, above. |
| Web opt-out — Google-Extended | Exists and is documented, and three properties limit it, all from Google's own text: it is opt-out (silence means yes); it is prospective — it governs "training future generations of Gemini models", so material already in the corpus is unaffected; and it is site-owner-only, settable by whoever controls robots.txt rather than by the author of the words. Google's own description of what it does not cover is "some of Google's other systems", and it publishes no enumeration of which. developers.google.com/crawling, 1 Sep 2026 |
| 🔴 What could not be established | Which model wrote any given piece. Training data for the models actually in the app. Whether a paid consumer subscription changes training — docs silent across 11 enumerated sections. What "not used directly for training" means for Notebook sources. Which systems Google-Extended does not cover. The scope of the separate Google-GeminiNotebook token. Who the human reviewers are. Which region serves a request — "This data may be stored transiently or cached in any country in which Google or its agents maintain facilities." |
| Sources | support.google.com/gemini/answer/13594961 · answer/13278892 · ai.google.dev/gemini-api/terms · Workspace Privacy Hub · developers.google.com/crawling · deepmind.google model cards · Google 2026 Environmental Report · Google AI-at-scale environmental paper — all 1 Sep 2026 |
6 · ChatGPT — OpenAI
| Field | Record |
|---|---|
| Tool, vendor, model/version | ChatGPT, OpenAI. GPT-5.6, in variants Sol / Sol Pro / Terra / Luna. "GPT-5.6 Luna is becoming the default model for Free and Go users"; "GPT-5.6 Sol powers Instant, Medium, High, and Extra High on eligible paid plans". ⚠ Defaults rotate without notice — three changes are visible in the release notes between March and August 2026. 1 Sep 2026 |
| What I use it for | Writing — scripts and prose |
| Training data — what IS published | The system card has a section called "Model Data and Training", and it says: "Like OpenAI's other models, GPT-5.6 was trained on diverse datasets, including information that is publicly available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate." No dataset list, no source inventory, no proportions. ⚠ The GPT-4 Technical Report's 2023 refusal — "this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar" — was never revoked. The 2026 system card omits the same categories without stating a reason, so that 2023 paragraph is the only published reason on record. 1 Sep 2026 |
| Consent basis | undisclosed. Named licensing deals exist (Axel Springer explicitly covers training; News Corp's announcement does not mention training) but no document states what fraction of the corpus is licensed, and the help text names "publicly available internet content" as a separate, first-listed bucket. Crawling is admitted; composition is not. |
| 🔴 Does it train on my content? | Yes, on by default for individuals. "When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models." The opt-out: "You can opt out of training through our privacy portal by clicking on 'do not train on my content.'" Business, Enterprise and API are excluded by default. Uploads are treated identically to typed text — "Content" is defined to include "files, images, audio and video", and no document draws a distinction. Opting out of training is not opting out of retention. 1 Sep 2026 |
| Data labour — the one place in this stack with actual numbers, and they are not the vendor's | Vendor: docs silent. Enumerated the GPT-5.6 System Card (9 sections), the foundation-models help article (9), "Our approach to data and AI", Terms (14), Enterprise Privacy (7). The only genuine vendor disclosure is four years old and covers a superseded model: the InstructGPT paper, "we hired a team of about 40 contractors on Upwork and through ScaleAI" — and it gives no pay rates. Journalism, and it is not vendor disclosure: TIME (Billy Perrigo, 18 Jan 2023) documented OpenAI's use of Sama, employing workers in Kenya, Uganda and India, taking home "between around $1.32 and $2 per hour depending on seniority and performance", labelling text "describing situations in graphic detail like child sexual abuse, bestiality, murder, suicide, torture, self harm, and incest". Kenyan moderators later petitioned Parliament. OpenAI's response: "we establish and share our own ethical and wellness standards for our data annotators" — ⚠ that standard could not be located as a published document. |
| Environmental — all five metrics | Nothing published. No sustainability, ESG, environmental or climate report or page exists on openai.com; openai.com/sustainability returns 404. WUE, PUE and water source: absent from every document enumerated (Abilene post, Project Camellia post, Stargate Norway post, OSTP submission, GPT-5.6 system card). Heat reuse: one qualitative single-site commitment at Stargate Norway, no quantum, no offtaker. Renewables: Norway "will run entirely on renewable power" — no matching basis stated, hourly or annual — and no corporate-wide figure exists. ⚠ Sam Altman's "0.34 watt-hours" per query is a personal-blog assertion with no methodology, never substantiated in any OpenAI document. It should not be quoted as an OpenAI figure. 1 Sep 2026 |
| The chain, and where it breaks | It breaks at the first link. OpenAI names Microsoft, AWS and Oracle as compute providers and names datacentre sites — but region selection is an Enterprise/Edu feature only, in three regions, and even there "does not confine all processing to the region". For a consumer plan, OpenAI discloses nothing about which provider, which region or which site serves a request. Because the provider itself cannot be identified, no provider's figures can be attributed — not even as a range. |
| 🔴 What could not be established | The training corpus of any GPT-5.6 variant. The licensed / crawled / user / synthetic proportions. Whether any specific work is in the corpus — no vendor mechanism exists to ask. Whether the News Corp deal covers training. Whether blocking GPTBot affects already-trained data. 🔴 Whether Media Manager — the rightsholder opt-out tool OpenAI promised — exists in any form. The page still reads "Our goal is to have the tool in place by 2025", unrevised, twenty months past its deadline. Architecture, parameter count and training compute. Any annotation vendor used after Sama. Current annotator pay, hours or geography. The "ethical and wellness standards" OpenAI says it shares. All five environmental metrics. GHG emissions of any scope. |
| Sources | help.openai.com/en/articles/20001354 · deploymentsafety.openai.com GPT-5.6 System Card · arXiv 2303.08774 · help.openai.com/en/articles/7842364 · openai.com/index/approach-to-data-and-ai · openai.com/policies/terms-of-use · help.openai.com/en/articles/5722486 · openai.com/enterprise-privacy · developers.openai.com/api/docs/bots · help.openai.com/en/articles/9903489 · openai.com Stargate/Abilene/Camellia posts · OpenAI OSTP RFI · TIME 18 Jan 2023 · ITWeb Africa 18 Jul 2023 · TechCrunch 1 Jan 2025 — all 1 Sep 2026 |
7 · Claude — Anthropic
Disclosure about the enquiry itself: this row was researched using Claude, examining Anthropic. The standard applied was deliberately stricter rather than gentler, and the three findings least flattering to Anthropic are stated first below.
| Field | Record |
|---|---|
| Tool, vendor, model/version | Claude, Anthropic PBC. Claude 5 generation: claude-opus-5, claude-sonnet-5, claude-fable-5, plus claude-haiku-4-5. 1 Sep 2026 |
| What I use it for | Writing — scripts and prose; and this enquiry |
| Training data — what IS published | Six sentences, and that is the whole disclosure: "Claude Opus 5 was trained on a proprietary mix of publicly available information from the internet, public and private datasets, and synthetic data generated by other models… We use a general-purpose web crawler called ClaudeBot to obtain training data from public websites. This crawler adheres to industry-standard practices with respect to the 'robots.txt' instructions… We do not access password-protected pages or those that require sign-in or CAPTCHA verification. We conduct due diligence on the training data that we use." Claude Opus 5 System Card §1.1, 24 Jul 2026, checked 1 Sep 2026 |
| 🔴 Contradiction — Anthropic's two official descriptions of the same model do not match | System Card: "…publicly available information from the internet, public and private datasets, and synthetic data generated by other models." Transparency Hub, same model: "…publicly available information from online sources, public and private datasets, user data, and synthetic data generated by other models." Both live on 1 Sep 2026. Neither is marked as superseding the other. They differ on whether user data is a training input for the flagship model. |
| 🔴 Disclosure went backwards between versions | Claude Opus 4.5's entry named "data provided by data-labeling services and paid contractors", "data from Claude users who have opted in", and a date bound ("up to May 2025"). Every 5-series entry drops all three, replacing them with "public and private datasets". Same page, same field, newer model, less disclosed. anthropic.com/transparency/model-report, 1 Sep 2026 |
| Consent basis | scraped — Anthropic states in its own system card that it uses ClaudeBot to obtain training data from public websites. The mechanism is admitted; the composition is withheld. Not licensed corpus: Anthropic told ADWEEK (26 Aug 2026) it "does have data partnerships" and "declined to share further specifics"; it has announced none publicly and scored 0 on Stanford CRFM's licensed-data-sources indicator by declining to name its top five. |
| 🔴 Does it train on my content? | Per the binding terms, yes unless I opt out — and Anthropic's own documents disagree about this. The announcement says "giving users the choice to allow"; the Consumer Terms and Privacy Policy both say "unless you opt out through your account settings"; and the settings help article — the one place a user would look — does not state the default either way. Two carve-outs survive opting out: feedback, and "your Materials are flagged for safety review". ⚠ That second one matters here specifically. I write political satire. Material touching political, violent or sexual subject matter is exactly what a safety classifier flags — so it is the carve-out most likely to catch my actual work. Turning it off is not retroactive: "Your data will still be included in model training that has already started and in models that have already been trained." Retention with training on: 5 years. Off: 30 days. Team, Enterprise and API are not trained on. |
| Data labour | docs silent, with one paragraph of acknowledgement: "Anthropic partners with data work platforms to engage workers… Anthropic will only work with platforms that are aligned with our belief in providing fair and ethical compensation to workers… following our crowd worker wellness standards detailed in our procurement contracts." No vendor named. No wage. No country. No worker count. No hours. ⚠ The "crowd worker wellness standards" are referenced and not published — searched by exact phrase; enumerated the legal index (11 documents) and the full site footer (9 categories). Stanford CRFM scored Anthropic 0 on data-laborer practices, noting no compensation information and no identification of worker countries. There is no named investigative journalism and no peer-reviewed work on Anthropic's annotation workforce. 1 Sep 2026 |
| 🔴 Environmental — WUE, PUE, water source, heat reuse, renewables | All five: declined in writing. anthropic.com/sustainability returns HTTP 404. Enumerated with no environmental content: the company page and full site footer, the legal index (11 documents), the Transparency Hub (45 section headings, 28 links), and both current system cards. Asked by Stanford CRFM for its figures, Anthropic answered — on energy usage, carbon emissions and water usage alike — "This information is proprietary and not disclosed publicly to protect competitive advantages and intellectual property." No SBTi target. No B-Corp certification. No published public benefit report of any year. No clean power purchase agreement announced; not a member of the Corporate Energy Buyers Association. It does buy carbon removal through Frontier Climate — which is not an emissions disclosure and not clean power, and tells us nothing about the five metrics. Stanford CRFM FMTI Dec 2025; Heatmap News 17 Jun 2026; checked 1 Sep 2026 |
| The chain, and where it breaks | It breaks one link earlier than everyone else's. Anthropic names its providers — AWS ("up to 5 gigawatts of capacity for training and deploying Claude"), Google Cloud/Broadcom ("multiple gigawatts of next-generation TPU capacity"), Microsoft Foundry — and names no datacentre site at all. Both compute announcements were enumerated for power source, renewables, water, PUE, WUE and carbon: neither contains a single environmental sentence. Water source is a site property; with no site named, that metric breaks before the region question arises. And since Anthropic makes no renewable claim of any kind, there is no claim to characterise as hourly or annual. Regional endpoints exist on Bedrock/Vertex/Foundry — a control, not a report — and the default global endpoint routes "to regions with available capacity" without disclosing where. |
| 🔴 What could not be established | Corpus size, date range, any named source, any named licensor. The crawled / purchased / synthetic split. A distinct training-data cutoff, which the docs promise and the Transparency Hub does not deliver. Whether the robots.txt opt-out affects already-trained models — the article is silent, enumerated in full. The default state of the training setting. Any data-labelling vendor named by Anthropic. The crowd worker wellness standards. Wages, hours, countries, headcount. All five environmental metrics. Any emissions disclosure of any scope. Any named Anthropic datacentre site. The Bartz primary court documents (the $1.5bn settlement received final approval 20 Jul 2026; it releases input-side claims through 25 Aug 2025 and explicitly does not release output-side or future-conduct claims — ⚠ sourced from a party-adjacent body, not the docket, and it is a settlement, not an admission). ⚠ trust.anthropic.com is JavaScript-rendered and could not be read — a gap in my checking, not a finding of silence. |
| Sources | Claude Opus 5 System Card · Claude Fable 5 & Mythos 5 System Card · anthropic.com/transparency/model-report · anthropic.com/legal/consumer-terms · /privacy · /commercial-terms · anthropic.com/news/updates-to-our-consumer-terms · privacy.claude.com ClaudeBot and retention articles · code.claude.com/docs/en/data-usage · claude.com/regional-compliance · anthropic.com/news compute announcements · anthropic.com/sustainability (404) · Stanford CRFM FMTI Dec 2025 · ADWEEK 26 Aug 2026 · Heatmap News 17 Jun 2026 — all 1 Sep 2026 |
Reachable, in the stack, not enquired into this year
Named for completeness under commitment 1. Neither is in current use, so neither was researched — researching a tool I am not using would imply a practice I do not have.
- ElevenLabs — voice backup. Live account, remediation path only if a prompted accent fails. Not yet needed in production. Not enquired into.
- HeyGen — free account, unupgraded, unused for seven months. Named fallback if the lip-sync path fails. Not enquired into.
If either enters use, it gets the full enquiry — either at the next annual review or out of cycle, whichever comes first.
Tier B — named, not enquired into
Tools that do not generate work and were not trained to produce it. Commitment 1 names them; commitment 2 does not bite.
Filmora (editing) · Canva (design) · Krystal.io (hosting, DNS) · Google Drive (storage) · Proton Mail (email) · Airtable (prompt library) · Buffer (scheduling) · YouTube and Substack (publishing).
How to check any of this
Every claim above carries a URL and the date I checked it. Open the link, find the quote. Where a cell says docs silent, the document I read is named — open it and see that the thing is not there.
Three documents defeated my tools and should be opened by hand before the next review, because they may contain figures I have recorded as "not established":
- Google's 2026 Environmental Report, environmental data appendix (~pp. 82–97) — would settle whether Google publishes a fleet water-efficiency figure. The report is 112 pages; I could retrieve only about the first 38.
- Kuaishou's 2025 ESG Report, Key Environmental Indicators appendix — would settle group water consumption.
- trust.anthropic.com — the only Anthropic page I could not read.
One correction I made to my own work, recorded because the standards require it. During verification I found that a widely-repeated version of Google's per-prompt figure — presenting the energy, carbon and water numbers as one continuous quotation — does not exist in Google's paper. It is two separate sentences from different sections, with two verbs changed. I have quoted them separately above. Had that not been caught, a fabricated quotation would have appeared on a page whose entire premise is that I make no false claim.
Enquiry run 1 September 2026. Evidence base: 2026-09-01-standards-enquiry.md. Reviewed annually — next review September 2027, and I will publish the outcome either way, including if nothing has changed.