[This article is the significantly updated version of the following preprint: Taravinyas: Mapping Paninian Phonology to Digital Logic for Continuous Touchscreen Input. One can read it in order to see how I treated the Boolean expressions in the Brahmi script.]
It was a really fine evening. I woke up from my afternoon slumber, half-slept. A month before, I already had the eccentric idea of making a better keyboard for every Indic lanugage, based on the 12-key Kana keyboard.
I imagined a boolean circuit, where the validity lines of aspiration, voicing, and nasals were there, and they made letters like this:
D2 = nasal
D1 = ~nasal & aspirated
D0 = ~nasal & voiced
This was based on the 5x5 grid of the North Indian languages.
Anyway, on the very small exercise book I had, after drawing the circuit, I posted that on Reddit, asking for people to join my research on making a mathematical model, because I myself didn't understand what I was daydreaming of. (Nobody accepted my invitation, though...)

And 10 months later, I have completed the model too.
Here it is: https://bandhanpramanik.github.io/tamil-input-engine-checkbox/
Try out the website at first, because the article is going to be very long.
This is a general approach I'm developing for Indic scripts. I started with a keyboard specifically for the Tamil script, a Dravidian script. More scripts coming soon.
Step 1: Get the table and the other things, all the caveats and stuff
This part will be heavy with linguistic jargons, but I will explain the real mathematical stuff from Step 2. For now, just look at the structure of the columns:
Tamil consonants: I have a table like this (self-made after referring to multiple sources):
| Places of Articulation | Vallinam | Mellinam | Idaiyinam | Sibilants (GRANTHA) |
|---|---|---|---|---|
| Velar | க் | ங் | - | - |
| Palatal | ச் | ஞ் | ய் | ஶ் |
| Retroflex | ட் | ண் | ழ் / ள் | ஷ் |
| Dental | த் | ந் | - | - |
| Labial | ப் | ம் | வ் | - |
| Alveolar | ற் | ன் | ர் / ல் | ஸ் |
There are three more Grantha characters taught from elementary school that we need to include: ஜ, ஹ, and க்ஷ.
Tamil vowels: Look at the table at first:
| Base vowel | Contrastive Monophthong | Diphthong |
|---|---|---|
| அ | ஆ | - |
| இ | ஈ | - |
| உ | ஊ | - |
| எ | ஏ | ஐ |
| ஒ | ஓ | ஔ |
Contrastive Monophthong simply means the one that is more marked. In Tamil Brahmi, people used to optionally put a dot, and that dot used to shorten the Tamil-Brahmi equivalent of ஏ into the Tamil-Brahmi equivalent of எ. The roles would have reversed then: the dotless ஏ would have become the base vowel, and எ the contrastive monophthong. However, in the Modern Tamil script, if you look at the letter itself, we have "எ" as unmarked, and "ஏ" as marked, which is why the table has been made this way.
Also, other characters include Pulli, Ayatam, and ZWNJ (to separate ligtaures, viz., க்ஷ into க்ஷ and ஸ்ரீ into ஸ்ரீ). These aren't vowels, but have been clubbed as extended characters.
Step 2: Create the coarse validity truth tables
Here's where the mathematical model starts. We will only talk about the consonants in this writeup.
From a row-major standpoint, "Coarse" literally means the rows, "Fine" means the columns. We will talk about Coarse positions as row number in binary, and Fine positions as the column number in binary.
Coarse validity means, given this Coarse position, which columns are even valid.
For example, look at the Velar row in the consonants. We do not have any Idaiyinam or Sibilants. However, in the Retroflex row, we have TWO Idaiyinam characters, and one Sibilant too. We even have rows where we have a single Idaiyinam character.
This means, that for Idaiyinam, we need something ternary: 0, 1, 2. However, we can use two boolean variables, and it will essentially be 00, 01, 10. For the validity of Sibilants, we can simply use 0 or 1.
I named the Idaiyinam Coarse Validity variable as C_I (C_I1 and C_I0) and C_S. C is the input variable here.
This is how the truth table looks like:
| C | C2 | C1 | C0 | C_I1 | C_I0 | C_S |
|---|---|---|---|---|---|---|
| க row | 0 | 0 | 0 | 0 | 0 | 0 |
| ச row | 0 | 0 | 1 | 0 | 1 | 1 |
| ட row | 0 | 1 | 0 | 1 | 0 | 1 |
| த row | 0 | 1 | 1 | 0 | 0 | 0 |
| ப row | 1 | 0 | 0 | 0 | 1 | 0 |
| ற் row | 1 | 0 | 1 | 1 | 0 | 1 |
We now have a function that goes from Coarse Position to Coarse Validity. These variables will be useful to know what places have letters, and where there's none.
Input: CoarsePosition
Output: CoarseValidity
Step 3: Create the Fine Positions Truth table
We saw CoarsePosition -> CoarseValidity, now we will see how FineFeatures -> FinePositions.
Now, what are FineFeatures? These are the specific columns allotted. Each column ticks a variable.
Input variables: Vallinam (V), Mellinam (M), Idaiyinam (I1 and I0), Sibilants (S)
Output variables: FinePosition (D2, D1, D0)
Here, for variable I, we have 3 values:
- Idaiyinam not selected: I1 = 0, I0 = 0
- The only idaiyinam letter selected / the left idaiyinam letter selected: I1 = 0, I0 = 1
- The right idaiyinam letter selected: I1 = 1, I0 = 1
Notice how the FineFeatures for Idaiyinam has 00, 01, 11, while the CoarseValidity for the same has 00, 01, 10. The reason is fully arbitrary and depends upon the convenience of the developer.
| D | D2 | D1 | D0 | S | I_1 | I_0 | M | V |
|---|---|---|---|---|---|---|---|---|
| ட | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| ண | 0 | 0 | 1 | 0 | 0 | 0 | 1 | 0 |
| ழ | 0 | 1 | 0 | 0 | 0 | 1 | 0 | 0 |
| ள | 0 | 1 | 1 | 0 | 1 | 1 | 0 | 0 |
| ஷ | 1 | 0 | 0 | 1 | 0 | 0 | 0 | 0 |
The best thing here is, according to your convenience, you can select which 6 binary numbers you want to use as FinePositions. I have taken 000 to 100, you can take something else.
Step 3: Find out the Boolean expressions using a Computer Algebra System
Now that we have the truth tables, here are the programs. Let's discuss one-by-one:
from sympy import symbols from sympy.logic import * from sympy.logic.boolalg import * c2, c1, c0 = symbols('c2,c1,c0') variables = [c2, c1, c0] # change this [ci1, ci0, cs]. note that c_i accepts 00, 01, and 10 only. i accepts 00, 01, 11 only. # each inner list corresponds with the respective rows in the lookup table values = [[0, 0, 1, 0, 0, 1], [0, 1, 0, 0, 1, 0], [0, 1, 1, 0, 0, 1]] ''' Explaining this part: I=11 and I=01 don't have any specfic meaning embedded by the Boolean expression. Their meanings are entirely set by the developer after looking at the letters present in those places. ''' length = len(variables) # change this input_ = [i for i in range(len(values[0]))] dontcare_minterms = list(set(range(2**length)) - set(input_)) dontcare_expr = SOPform(variables, dontcare_minterms) minterms = [] for i in range(length): # print() minterms.append([]) for j, ob1 in enumerate(values[i]): # print(f"{j:03b} | {ob1}") if ob1 == 1: minterms[i].append(input_[j]) sop = [] for i in minterms: sop.append(SOPform(variables, i)) print("C_I1 =", simplify_logic(sop[0], dontcare=dontcare_expr, form="dnf")) print("C_I0 =", simplify_logic(sop[1], form="dnf")) print("C_S =", to_anf(simplify_logic(sop[2], dontcare=dontcare_expr)))
Each row is inspected one at a time. We are actually converting the three lists of six CoarseValidity values into three lists of six MINTERMS. After that, we are converting that to the SOP form and minimizing it. And voila, we are done!
Code output:
C_I1 = (c0 & c2) | (c1 & ~c0)
C_I0 = (c0 & ~c1 & ~c2) | (c2 & ~c0 & ~c1)
C_S = c0 ^ c1
After simplification by hand, we have:
C_I1 = (c0 & c2) | (c1 & ~c0)
C_I0 = ~c1 & Xor(c0, c2)
C_S = c0 ^ c1
Now, for the Fine Positions thing, we have another program:
# Human-generated code; looked at the docs to make this from sympy import symbols from sympy.logic import * from sympy.logic.boolalg import * def find_minterms(output, minterms, length_output): output_bits = [] for i in range(length_output): output_bits.append([]) for j in output: output_bits[i].append((j >> i) & 1) minterms_d = [] for i in range(length_output): minterms_d.append([]) for j, ob1 in enumerate(output_bits[i]): if ob1 == 1: minterms_d[i].append(minterms[j]) return minterms_d # change this ci1,ci0, cs, s,i1,i0,m,v = symbols('ci1,ci0,cs,s,i1,i0,m,v') minterms = [1,2,4,12,16] variables = [s,i1,i0,m,v] dontcare_minterms = list(set(range(2**len(variables))) - set(minterms)) expr = SOPform(variables, minterms) dontcare_expr = SOPform(variables, dontcare_minterms) # change this output = [i for i in range(len(minterms))] length = len(minterms).bit_length() minterms_d = find_minterms(output, minterms, length) sop_d = [] for i in minterms_d: sop_d.append(SOPform(variables, i)) for i, final_expr in enumerate(sop_d): print("D" + str(i) + " =", simplify_logic(final_expr, dontcare=dontcare_expr)) # change these lines to include what's allowed coarse = expr & (~i0 | (ci1 ^ ci0)) & (~i1 | (i0 & (ci1 & ~ci0))) coarse &= ~(ci1 & ci0) & (~s | cs) & Exclusive(s, i0, m, v) invalid_all = Not(coarse) invalid_all = simplify_logic(invalid_all, force=True) print("Invalid =", invalid_all)
This program can be divided into two phases:
- Finding FinePositions from FineFeatures, strictly based on the truth table, exactly as we did earlier.
- Finding the exact things that are invalid.
Let's talk about Number 2.
At first, expr finds out everything that's valid.
In the variable called coarse, there are certain rules in order to make our input valid:
i0 -> (ci1 ^ ci0)
i1 -> (i0 & (ci1 & ~ci0))
s -> cs
Then, by the rule that a -> b can we written as (~a | b), we type out the expressions.
These rules are clubbed with two basic requirements such as C_I not being 11 and only one of S, I_0, M, V being 1.
The reason we are using CoarseValidity variables here is because we need to know if the place is valid or not in the first place.
We are using the conjunctive normal form (CNF) for combining different rules.
Now, after finding out everything valid, we take the NOT of that expression, and we print its simplified form. This will act as the validity predicate.
That's it for the boolean expressions.
Step 4: Create the evaluation method
I used Typescript for this.
For CoarsePositions -> CoarseValidity, I made a method. For FineFeatures -> FinePositions, I made another method. For Invalid_D, another method was made.
However, remember the three additional Grantha consonants we talked about? Well, they are the inhabitants of the ExtendedWorld:
interface NormalWorld { features: FineFeatures, validity: CoarseValidity } interface ExtendedWorld { e1: boolean, e0: boolean } type World = NormalWorld | ExtendedWorld;
As we stated, ternary values require two bits to express. That's exactly what e1 and e0 serve as.
Here's the exact evaluation function:
const FLAG_3: number = 0b1000; const FLAG_2: number = 0b0100; const FLAG_1: number = 0b0010; const FLAG_0: number = 0b0001; export function evalD(alpha: boolean, world: World): number { if (alpha && "e0" in world) { return FLAG_3 | (world.e1 ? FLAG_1 : 0) | (world.e0 ? FLAG_0 : 0); } else if (!alpha && "features" in world) { const invalid_d = findInvalidD(world.validity, world.features); if (invalid_d) return 0b0111; const abc = findFinePositions(world.features); return (abc.d_2 ? FLAG_2 : 0) | (abc.d_1 ? FLAG_1 : 0) | (abc.d_0 ? FLAG_0 : 0); } else return 0b0111; }
alpha = 0 meant the normal mode, for which we made the table. alpha = 1 is for the extended mode. As we stated, invalid_d acts as a validity predicate.
Because this framework is fundamentally suitable even for the embedded systems, we went for an absurdly optimized way of writing TS, just to demonstrate how optimized it can be made.
The output will be a number, and we will actually look at what the number is.
Step 5: Make the lookup table
I heard that in TypeScript, an array-based lookup table takes the least amount of time, so I am using it.
This is my table:
type StrOrUndef = string | undefined; type SixConsonantGroups = [StrOrUndef, StrOrUndef, StrOrUndef, StrOrUndef, StrOrUndef, StrOrUndef]; const CONSONANT_LOOKUP_TABLE: Array<SixConsonantGroups | string | undefined> = [ /*0b0000*/ ["க", "ச", "ட", "த", "ப", "ற"], /*0b0001*/ ["ங", "ஞ", "ண", "ந", "ம", "ன"], /*0b0010*/ [undefined, "ய", "ழ", undefined, "வ", "ர"], /*0b0011*/ [undefined, undefined, "ள", undefined, undefined, "ல"], /*0b0100*/ [undefined, "ஶ", "ஷ", undefined, undefined, "ஸ"], /*0b0101*/ undefined, /*0b0110*/ undefined, /*0b0111*/ "INVALID", /*0b1000*/ "ஜ", /*0b1001*/ "ஹ", /*0b1010*/ "க்ஷ", /*0b1011*/ undefined, /*0b1100*/ undefined, /*0b1101*/ undefined, /*0b1110*/ undefined, /*0b1111*/ undefined, ];
Notice how changing the outputs for the FinePositions matter, how important it is to keep proper values for the least significant three 3 bits. That's why I kept the sympy program so flexible.
Also notice how a properly written program will never return undefined from the lookup table.
Step 6: Just... document how the inputs should be given.
For CoarsePositions,
Velar: 000
Palatal: 001,
...
Alveolar: 101
For FineFeatures,
Vallinam: 00001
Mellinam: 00010
Idaiyinam (Normal/Rhotic-like): 00100
Idaiyinam (Lateral): 01100
[Grantha Mode] Sibilant: 10000
And in the extended mode, the only valid values are 00, 01, and 10. 11 shouldn't be inputted. (I guess I'll have to make that explicit in my code...)
Thanks guys.
This is it. This is the entire mathematical framework I made for Tamil, in order to make the structure of the writing system as the structure of the keyboard, just like the Japanese 12-key keyboards do.
Check out the sympy code for the Coarse truth table: https://github.com/BandhanPramanik/taravinyas-sympy/blob/main/tamil/tamil-coarse-consonants.py
Check out the sympy code for the Fine truth table: https://github.com/BandhanPramanik/taravinyas-sympy/blob/main/tamil/tamil-consonants.py
Check out the typescript code: https://github.com/BandhanPramanik/tamil-input-engine-checkbox/blob/main/ts/consonant.ts