davejonescue
Puritan Board Graduate
So today I was thinking something, as I am trying to come up with ideas possibly of getting closer to reading fluency in as little as time as possible. Of what I understand it, while grammar rules are important, so is building vocabulary. This is what got me thinking, when attempting to transcribe the text for translation, the hard part is cleaning the facsimiles and looking for mess-ups. But, what if instead of trying to produce translations, the goal is to transcribe simply to get a list of unique words and word frequencies? So, say I transcribe what are known as the "pivotal" Reformed Scholastics, (if I can find a list), generate a word list; and then, turn those words into phonetic renderings as to be able to create audio flash cards to listen to, on the words (starting from the most used, to the least) from those authors?
There are a few reasons why this is possible:
1. Claude is at 99% accuracy.
2. Most of the matching, phonetic conversion, frequency counts can be automated.
3. This will potentially get folks prepared to read "the majors" by knowing the vocab going in.
Someone on here said that there are only so many 1000's of words one has to have down to get going really good. I can't find that list. I was able to get all the words and definitions from something called Latdic, (which has a public domain notice) but that consists of about 40,000 words, which seems impossible to memorize (and many are probably not needed). I don't think it would be hard to compare the transcripts list to that list, see what there are already definitions for (as to create either digital/audio flashcards) and find the definitions to the words that are not there.
Just seems like a way to take the present tech capabilities and give a little push in the right direction to being able to read these books fluently when/if someone ever attempts to read them.
What do you guys think?
There are a few reasons why this is possible:
1. Claude is at 99% accuracy.
2. Most of the matching, phonetic conversion, frequency counts can be automated.
3. This will potentially get folks prepared to read "the majors" by knowing the vocab going in.
Someone on here said that there are only so many 1000's of words one has to have down to get going really good. I can't find that list. I was able to get all the words and definitions from something called Latdic, (which has a public domain notice) but that consists of about 40,000 words, which seems impossible to memorize (and many are probably not needed). I don't think it would be hard to compare the transcripts list to that list, see what there are already definitions for (as to create either digital/audio flashcards) and find the definitions to the words that are not there.
Just seems like a way to take the present tech capabilities and give a little push in the right direction to being able to read these books fluently when/if someone ever attempts to read them.
What do you guys think?
