# The Rigveda by stratum: a corpus analysis

**Method note first, because everything below depends on it.**

Corpus: VedaWeb (Cologne Center for eHumanities), full Rigveda, TEI-XML with the Zurich
lemmatisation and morphology. **164,758 lemmatised tokens, 10,031 unique lemmas.**
Licence CC-BY 4.0.

Chronology: **Arnold's metrical strata** (*Vedic Metre in its Historical Development*, 1905),
compiled into machine-readable form by Dieter Gunkel (Richmond) and Kevin M. Ryan (Harvard).
Every *line* carries a stratum label: **A** (Archaic), **S** (Strophic), **C** (Cretic),
**N** (Normal), **P** (Popular). This is finer-grained than mandala number, and independent
of it.

Validation that the labels behave chronologically: Popular is **46.7% of Book 10** and
2.6–6.2% of the family books; Archaic peaks in **Book 6 (62.6%)**, which is Witzel's oldest.
The two ends, A and P, are trustworthy. The middle ordering (S/C/N) is Arnold's and is not
independently verified here — **all claims below use the A-vs-P contrast only.**

Semantic categories are **my hand-curation** and are the softest part of this. They are
printed in full so they can be argued with.

---

## 1. The headline test: does retroflexion increase over time?

The substrate hypothesis predicts that retroflex consonants — absent from reconstructed PIE,
pervasive in Dravidian — entered Indo-Aryan through contact, and should therefore be
**rarer in the oldest layer**.

The diagnostic is not raw retroflexion. Sanskrit generates retroflexes internally by regular
rule (RUKI: ṣ after r/ṛ/u/i/k; nati: n→ṇ after r/ṛ/ṣ). What matters is retroflexion that
**those rules do not explain**. My filter for this is crude — it checks whether any
conditioning environment precedes the retroflex in the word — and it will misclassify
compounds and sandhi forms. Treat the absolute numbers as approximate.

### Token level

| Stratum | Tokens | Any retroflex | **Unconditioned** |
|---|---|---|---|
| **A** Archaic | 33,930 | 8.48% | **0.825%** |
| S Strophic | 37,380 | 8.10% | 0.73% |
| C Cretic | 29,621 | 8.15% | 0.81% |
| N Normal | 40,896 | 8.44% | 0.98% |
| **P** Popular | 22,961 | 7.93% | **1.215%** |

A vs P: z = 4.63, p < 0.001.

### CORRECTION (added after the fact) — this result does not survive

The `lemma` field is a **Grassmann dictionary form**: normalised, sandhi undone, retroflexion
regularised by the lexicographers. Re-running the identical test on the `surface` field — the
actually attested word forms — gives:

| Field | A (Archaic) | P (Popular) | z |
|---|---|---|---|
| lemma (dictionary form) | 0.825% | 1.215% | 4.63 |
| **surface (attested form)** | **1.444%** | **1.476%** | **0.32** |

**On the transmitted text the effect vanishes.** The token-level trend was a property of
editorial normalisation, not of the Rigveda. The type-level null below already pointed at
this; the surface test identifies the mechanism.

There is a further problem the test cannot escape: the Rigveda was fixed orally and
retroflexion across word boundaries is partly imposed by the śākala recension and the
padapāṭha analysis. **Even the surface forms are the text as the tradition standardised it.**
A diachronic phonological count on a corpus this heavily normalised may not be a well-posed
test at all — which is why Kuiper and Witzel argue from individual etymologies instead.

### Type level — and this is where it falls apart

| Stratum | Unique lemmas | Unconditioned-retroflex types |
|---|---|---|
| A | 4,448 | **2.00%** |
| S | 4,369 | 1.72% |
| C | 4,144 | 1.83% |
| N | 4,537 | 1.87% |
| P | 3,743 | **2.00%** |

**Identical at both ends.** The token-level rise is therefore driven by a small number of
high-frequency words repeating more often in late hymns — *not* by the late strata drawing on
a broader foreign-looking vocabulary.

**Verdict: the claim fails on two independent grounds — flat at type level, and absent
entirely on surface forms.** A significant
token-level trend with a flat type-level distribution is the signature of a frequency effect,
not a lexical-influx effect. Anyone quoting the first table without the second is
overstating the evidence.

What the result does *not* rule out: that specific borrowed items entered late (see §3),
or that the substrate entered before the Rigveda's oldest stratum and is therefore invisible
to this test. The second possibility is the one Kuiper and Witzel actually argue, and this
analysis cannot address it.

---

## 2. Rudra — the correction

I had called Rudra marginal. **The corpus says otherwise.**

| Deity | n | Family books 2–7 | Book 10 | A/S/C/N/P |
|---|---|---|---|---|
| índra | 2,435 | 972 | 335 | A23 S30 C14 N24 P09 |
| agní | 1,724 | 880 | 314 | A22 S21 C20 N27 P10 |
| dyú- ~ div- | 1,008 | 393 | 157 | A18 S22 C23 N26 P11 |
| sóma | 977 | 230 | 113 | A13 S21 C11 N44 P11 |
| gáv- ~ gó- | 543 | 193 | 91 | A19 S16 C18 N33 P14 |
| váruṇa | 396 | 200 | 55 | A18 S27 C22 N24 P09 |
| **rudrá** | **135** | **69** | 22 | **A23 S14 C27 N28 P07** |
| bŕ̥haspáti | 123 | 52 | 51 | A01 S16 C33 N29 P21 |
| **víṣṇu** | **98** | 39 | 13 | A22 S34 C10 N20 P13 |

**Rudra is more frequent in the Rigveda than Viṣṇu**, is well represented in the oldest
stratum (A23), and is concentrated in the family books. He is not a late arrival and not a
demonised outsider. He is a genuinely old Vedic god who is *feared and propitiated* rather
than celebrated — which is a different thing from being demonised.

`śivá-` occurs **51 times, always as an adjective** ("auspicious, kindly"), evenly spread
across strata (A27 S18 C18 N18 P20). Never a name. The euphemism is Vedic; the deity called
by it is not.

**The genuinely damning fact for a Vedic origin of Śiva-worship is not Rudra's frequency —
it is `śiśnádeva-`.** Twice in the Rigveda, both times an insult. The later cult's central
form of worship is the thing the Rigveda mocks.

---

## 3. Absences — verified, and the ones I got wrong

I initially reported several absences that were my own lemma-lookup failures. Separating them:

### Verified absent from the entire Rigveda
| Word | Meaning | Why it matters |
|---|---|---|
| **vrīhí** | rice | The staple of later South Asia. Absent. Rigvedic grain is `yáva`, barley (n=23). |
| **godhū́ma** | wheat | Absent. |
| **vyāghrá** | tiger | **Absent.** The tiger is on Indus seals in quantity. A text composed in the tiger's range with no word for it is a geographic argument. |
| **kulāla** | potter | Absent. |
| **vṛścika** | scorpion | Absent. |
| **dhánus** | bow | Absent as a lemma — the weapon appears under other terms; `íṣu-` (arrow) n=11, and **55% of those are Popular stratum**. |
| **sárasvatī** (bare) | river/goddess | Present only in derivatives (`sárasvatīvant-`, `sārasvatá-`) in this lemmatisation. |

### My errors — these ARE in the corpus
`mayūrī́-` and `mayū́raroman-` (peacock, 1 each), `naú- ~ nā́v-` (boat, n=40),
`vŕ̥ka-` (wolf, n=30), `gŕ̥dhra-` (vulture, n=8), `bŕ̥haspáti-` (n=123),
`dyú- ~ div-` (n=1,008), `gáv- ~ gó-` (n=543), `vŕ̥ṣan-` (n=458).

Corrections published, per house rule.

---

## 4. The outsider vocabulary is tiny, and it is front-loaded

| Term | n | Family 2–7 | Bk 10 | A/S/C/N/P |
|---|---|---|---|---|
| dásyu | 67 | 31 | 8 | A27 S25 C18 N27 **P03** |
| ásura | 71 | 29 | 19 | A17 S20 C23 N23 P18 |
| paṇí | 50 | 22 | 11 | A26 S14 C22 N18 P20 · **unconditioned ṇ** |
| dā́sa | 23 | 10 | 11 | **A39** S13 C30 N04 P13 |
| rákṣas | 80 | 33 | 15 | A18 S18 C08 N26 **P31** |
| yātudhā́na | 25 | 4 | 20 | **P96** |
| śiśnádeva | **2** | 1 | 1 | A50 S50 |
| śímyu | **2** | 1 | 0 | C100 |
| yákṣu | **2** | 2 | 0 | C100 |
| kī́kaṭa | **1** | 1 | 0 | C100 |
| **anā́s** | **1** | 1 | 0 | N100 |
| piśā́ci | **1** | 0 | 0 | P100 |
| vaikarṇá | **1** | 1 | 0 | C100 |
| śígru | **1** | 1 | 0 | C100 |

**`anā́s` — the "noseless"/"speechless" word that anchors a century and a half of racial
argument about the Rigveda — occurs exactly once.** So do kī́kaṭa, piśā́ci, vaikarṇá and
śígru. `śiśnádeva` occurs twice.

And the pattern in the frequent terms runs the *opposite* way to a hardening-racism model:
**`dásyu` and `dā́sa` are Archaic-weighted and fade** (dásyu is 3% Popular; dā́sa is 39%
Archaic). What *rises* into the late strata is demonological, not ethnic: `rákṣas` 31% Popular,
`yātudhā́na` **96% Popular**. The enemy shifts from a people to a demon.

For contrast: **`ā́rya-` occurs 38 times.** `jána-` (people) occurs 327 times.

---

## 5. Late-stratum vocabulary: where the borrowed-looking words actually sit

Items occurring wholly or almost wholly in the **Popular** stratum:

| Lemma | Gloss | n | Stratum |
|---|---|---|---|
| **lā́ṅgala** | plough | 1 | **P100** |
| **sī́tā** | furrow | 2 | **P100** |
| maṇḍū́ka | frog | 8 | **P100** · unconditioned ṇḍ |
| kapóta | dove | 6 | P83 · ka- prefix |
| úlūka | owl | 1 | P100 |
| khadirá | acacia | 1 | P100 |
| śalmalí, śimbalá | silk-cotton tree | 1 each | P100 |
| óṣadhi | herb | 8 | P75 · unconditioned ṣ |
| yātudhā́na | sorcerer | 25 | P96 |
| śraddhā́ | faith | 19 | P79 |
| yamá | Yama | 66 | P65 |
| sárva | (Śarva) | 68 | P72 |

**`lā́ṅgala` (plough) and `sī́tā` (furrow) are the two classic proposed non-Indo-European
agricultural loans, and both occur only in the latest stratum, one to two times each.**
That is the single most suggestive result in this analysis — a *targeted* late entry of
foreign agricultural vocabulary, which is exactly what §1's type-level null result says is
NOT happening at the level of the whole lexicon.

Both things are true at once: no general late influx, and specific late loans in
agriculture and local fauna and flora.

`úṣṭra` (camel, n=5) sits at **A40 S40** — old, and per Lubotsky belongs to the
**Central Asian / BMAC substrate shared with Iranian**, not the South Asian one.
Two different foreign layers, and this analysis does not separate them.

---

## 6. Śiva: what the etymology can and cannot settle

**Mainstream Indo-European derivation.** `śivá-` "auspicious, kindly, dear" ← PIE
**\*ḱeiwo-** "dear, one's own, familiar" (compare Latin *cīvis* "citizen", Old English *hīw*
"household"), from the root \*ḱei- "to lie, settle, be familiar". Semantically this is a
**euphemism**: the terrible one is addressed as "the friendly one" so as not to provoke him.
That is a well-attested Indo-European taboo-naming habit and it fits the Rigvedic usage
exactly — 51 occurrences, all adjectival.

**Rudra** ← √rud "to howl, cry" — "the Howler". A competing derivation links it to
\*h₁reudh- "red" (cf. *rudhira*, blood). Both are proposed; neither is settled.

**The Dravidian alternative you're pointing at is real and phonologically plausible.**
Dravidian \*cem/\*civ- "red" (Tamil *civappu* red, *cem-* red) would give *civa*. Śiva is
associated with red; Rudra may mean "red". So there are two homophonous candidate roots
converging on the same deity.

**Verdict, stated honestly:** the IE derivation is the mainstream one and is well formed.
The Dravidian derivation is not accepted in mainstream comparative linguistics, and I have
not seen it defended with a regular sound-correspondence argument rather than a resemblance
argument. **But etymology cannot settle your actual question anyway.** A Sanskrit word can
be applied to a non-Sanskrit god — that is precisely what "Murugan = Skanda" is. The name
being Indo-European tells us about the *word*, not about the *worship*. And on the worship,
the corpus evidence is unambiguous: `śiśnádeva` is an insult in the Rigveda.

### Your phonological diagnostic — the answer is "sometimes, and here is why not always"
You asked whether an indigenous word would show *l* for *r*, *v* for *y/b*, etc. Partly:
- **l/r**: yes, this is a real diagnostic axis, but it points **east**, not south — the
  Śatapatha mocks *he 'lavo* for *he 'arayo*, and eastern Middle Indo-Aryan (Māgadhī)
  systematically has *l* for *r* (Ashokan *lāja* for *rāja*).
- **Retroflexes**: the strongest diagnostic, tested in §1.
- **Initial ka-/ki-/ku-**: Kuiper's proposed prefixing substrate ("Para-Munda"). Flagged in
  the tables (`kapóta`, `karambhá`, `kúbhā`).
- **Geminates and word-final clusters**: secondary diagnostics, flagged.

The limit: these are *probabilistic* flags on individual words, not a decision procedure.
No single feature proves borrowing, and Sanskrit's own rules generate several of them
natively — which is exactly the confound that sank the naive version of §1.

---

## 7. What this analysis cannot do

- **It cannot separate the two substrate layers.** Lubotsky's Central Asian/BMAC layer
  (shared with Iranian, e.g. `úṣṭra`) and a South Asian layer are conflated here. Separating
  them requires an Avestan cognate check on every flagged lemma.
- **It has no Dravidian or Munda comparanda built in.** Every "possible loan" above is
  flagged on Sanskrit-internal phonological grounds only. A real test requires a Dravidian
  etymological database (DEDR) joined to the lemma list.
- **It covers only the Rigveda.** The other three Samhitas are not in this corpus dump. Your
  per-Veda comparison — where retroflexion, rice, iron and `mleccha` all appear — needs
  the Atharvaveda especially, since that is where the vocabulary changes most.
- **The semantic categories are mine.** Someone else's curation would give different totals.

## Reproducibility
`extract.py` (TEI → tokens TSV) · `lexicon.py` (categories) · `analyze.py` (counts,
diagnostics, significance). Corpus: github.com/VedaWebProject/vedaweb-data, CC-BY 4.0.
