View Single Post
Unread 04-13-2011, 05:33 PM   #21 (permalink)
THEINCREDIBLEdork
Emperor Meow
 
THEINCREDIBLEdork's Avatar
 

Join Date: Sep 2004
Posts: 9,316
Internets: 284585
THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute THEINCREDIBLEdork has a reputation beyond repute

Send a message via AIM to THEINCREDIBLEdork
Default

Quote:
Originally Posted by thekremlin View Post
I'm going to keep quoting "The Information" until someone cares. Apparently the english language is 75% redundant.

This is actually really cool. If you define information as an operation that eliminates uncertainty, an english sentance contains much less information than a string of letters. For instance if I write "nubblie", the following "s" doesn't really add any information. But in a string of letters where the letters all signify something, like a DNA code, every single letter adds the same amount of information. The author suggests that a random string of letters paradoxically contains more information than an english sentance, I'm still having mild trouble understanding how you can make that claim about random letters WITHOUT a prearranged code like DNA.

Apparently this is one of the ways they first started developing the AI that lets computers seem like they're having conversations. If the first letter of a word is "r", the probability that the next letter is an "x" is really low, whereas it's very probable that it's an "a", "e", or "o". "T" is followed by "h" more than any other letter. And so on. So you can write a program that makes strings of letters, and if you have the right weightings, you can make strings that kind of look like words. Now you go a step further, and do the same thing for 2-letter blocks, ie "ch" will never be directly followed by "su". This might seem obvious but you're "teaching" the program how to write words without telling it anything about the rules of english, just probabilities. And then you do it for words. Things like "if the word 'horse' is in a sentance, the word 'horse' is very likely to appear in that sentance again." And then you repeat that for 2-word codes, etc. Pretty soon (and by pretty soon I mean after years of collecting statistics and programming) you've got a computer program which can almost pass the Turing test, and you haven't actually told it any syntax or spelling rules.

That's fucking cool!!!!!!!

But yeah, Claude Shannon estimated that the english language was 75% redundant.
You spelled sentence incorrectly at least 4 times in that post. I do not think you know how to spell sentence.

#YOLO
THEINCREDIBLEdork is offline   Reply With Quote