Add Attribution & takedown section

#13
by madiedgar - opened
Files changed (1) hide show
  1. README.md +32 -2
README.md CHANGED
@@ -47,7 +47,7 @@ The hypothesis is **not** that non-English code matches or exceeds English code
47
 
48
  ## Base Model
49
 
50
- All adapters are trained on [CohereLabs/tiny-aya-base](https://huggingface.co/CohereLabs/tiny-aya-base) (3.35B parameters). Tiny Aya was chosen because it is small (deployable on a single 16 GB T4 GPU via QLoRA), accessible (Apache 2.0-licensed), and supports 70+ languages with explicit emphasis on lower-resourced ones — which makes the experimental ladder viable for `ur` at all.
51
 
52
  ## Adapter Inventory
53
 
@@ -222,4 +222,34 @@ Paper-grade evaluation results live on [`legesher/language-decoded-experiments`]
222
 
223
  ## License
224
 
225
- Apache 2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
47
 
48
  ## Base Model
49
 
50
+ All adapters are trained on [CohereLabs/tiny-aya-base](https://huggingface.co/CohereLabs/tiny-aya-base) (3.35B parameters). Tiny Aya was chosen because it is small (deployable on a single 16 GB T4 GPU via QLoRA), openly available (released under CC-BY-NC-4.0), and supports 70+ languages with explicit emphasis on lower-resourced ones — which makes the experimental ladder viable for `ur` at all.
51
 
52
  ## Adapter Inventory
53
 
 
222
 
223
  ## License
224
 
225
+ CC-BY-NC-4.0. The adapters inherit the license of the base model,
226
+ [CohereLabs/tiny-aya-base](https://huggingface.co/CohereLabs/tiny-aya-base)
227
+ (CC-BY-NC-4.0). The training datasets
228
+ ([legesher/language-decoded-data](https://huggingface.co/datasets/legesher/language-decoded-data))
229
+ are separately licensed under Apache-2.0.
230
+
231
+ ## Provenance, attribution & takedown
232
+
233
+ These adapters were fine-tuned from
234
+ [`CohereLabs/tiny-aya-base`](https://huggingface.co/CohereLabs/tiny-aya-base)
235
+ on a specific revision of the
236
+ [`legesher/language-decoded-data`](https://huggingface.co/datasets/legesher/language-decoded-data)
237
+ training conditions (see each adapter's configuration for the
238
+ condition and revision).
239
+
240
+ If you are the author of source code included in the training data
241
+ and would like attribution added or your code removed, open a
242
+ discussion on the dataset repository's **Community** tab or email
243
+ **support@legesher.com**. Removals are propagated in a new dataset
244
+ revision. Adapters already trained are frozen historical artifacts:
245
+ a dataset removal does not alter existing adapter weights, but we
246
+ will note affected conditions here and take reported concerns about
247
+ specific adapters into account.
248
+
249
+ **Usage caution.** These are research artifacts, not
250
+ production-ready models. Documented side effects include code
251
+ fragments leaking into natural-language output (strongest for Urdu
252
+ adapters) and matched-language regressions on specific evaluation
253
+ cells; the Condition 5 adapters were trained on corpora containing
254
+ raw translator output. Evaluate per language and per task before any
255
+ downstream use.