Baudot code

From Wikipedia, the free encyclopedia
Jump to: navigation, search

The Baudot code, invented by Émile Baudot,[1] is a character set predating EBCDIC and ASCII. It was the predecessor to the International Telegraph Alphabet No. 2 (ITA2), the teleprinter code in use until the advent of ASCII. Each character in the alphabet is represented by a series of bits, sent over a communication channel such as a telegraph wire or a radio signal. The symbol rate measurement is known as baud, and is derived from the same name.

History[edit]

Baudot code[edit]

Baudot code (Continental and UK versions)
Baudot code (Continental and UK versions)
Columns I, II, III, IV, and V show the code; the Let and Fig columns show the letters and numbers for the Continental and UK versions; The sort keys present the table in the order: alphabetical, Gray and UK
Europe sort keys UK sort keys
V IV I II III Con­ti­nen­tal Gray Let. Fig. V IV I II III UK
- - -
A 1 A 1
É & / 1/
E 2 E 2
I o I 3/
O 5 O 5
U 4 U 4
Y 3 Y 3
B 8 B 8
C 9 C 9
D 0 D 0
F f F 5/
G 7 G 7
H h H ¹
J 6 J 6
Figure Blank Fig. Bl.
Erasure Erasure * *
K ( K (
L = L =
M ) M )
N N £
P  % P +
Q / Q /
R R
S  ; S 7/
T  ! T ²
V ' V ¹
W  ? W  ?
X , X 9/
Z  : Z  :
t . .
Blank Letter Bl. Let.

Technically, five-bit codes began in the 16th century, when Francis Bacon developed the cipher now called Bacon's cipher. However, this cipher is not a machine cipher and as such is not readily suitable for telecommunications.[2]

Baudot invented his original code in 1870 and patented it in 1874. It was a 5-bit code, with equal on and off intervals, which allowed telegraph transmission of the Roman alphabet and punctuation and control signals. It was based on an earlier code developed by Carl Friedrich Gauss and Wilhelm Weber in 1834.[3][4][5] It was a Gray code (when vowels and consonants are sorted in their alphabetical order),[6] nonetheless, the code by itself was not patented (only the machine) because French patent law does not allow concepts to be patented.[7]

Baudot's original code was adapted to be sent from a manual keyboard, and no teleprinter equipment was ever constructed that used it in its original form.[8] The code was entered on a keyboard which had just five piano type keys, operated with two fingers of the left hand and three fingers of the right hand. Once the keys had been pressed they were locked down until mechanical contacts in a distributor unit passed over the sector connected to that particular keyboard, when the keyboard was unlocked ready for the next character to be entered, with an audible click (known as the "cadence signal") to warn the operator. Operators had to maintain a steady rhythm, and the usual speed of operation was 30 words per minute.[9]

The table on the right "shows the allocation of the Baudot code which was employed in the British Post Office for continental and inland services. It will be observed that a number of characters in the continental code are replaced by fractionals in the inland code. Code elements 1, 2 and 3 are transmitted by keys 1, 2 and 3, and these are operated by the first three fingers of the right hand. Code elements 4 and 5 are transmitted by keys 4 and 5, and these are operated by the first two fingers of the left hand."[8][10][11]

Baudot's code became known as International Telegraph Alphabet No. 1, and is no longer used.

Murray code[edit]

Paper tape with holes representing the "Baudot-Murray Code". Note the fully punched columns of "Delete/Letters select" codes at start of the message, these were used to cut the band easily between distinct messages. The message then starts by a Figure shift control followed by a carriage return.

In 1901, Baudot's code was modified by Donald Murray (1865–1945), prompted by his development of a typewriter-like keyboard. The Murray system employed an intermediate step, a keyboard perforator, which allowed an operator to punch a paper tape, and a tape transmitter for sending the message from the punched tape. At the receiving end of the line, a printing mechanism would print on a paper tape, and/or a reperforator could be used to make a perforated copy of the message.[12] As there was no longer a connection between the operator's hand movement and the bits transmitted, there was no concern about arranging the code to minimize operator fatigue, and instead Murray designed the code to minimize wear on the machinery, assigning the code combinations with the fewest punched holes to the most frequently used characters.[13][14]

The Murray code also introduced what became known as "format effectors" or "control characters" – the CR (Carriage Return) and LF (Line Feed) codes. A few of Baudot's codes moved to the positions where they have stayed ever since: the NULL or BLANK and the DEL code. NULL/BLANK was used as an idle code for when no messages were being sent, but the same code was used to encode the space separation between words. Sequences of DEL codes (fully punched columns) were used at start or end of messages or between them, allowing easy separation of distinct messages (BELL codes could be inserted in those sequences to signal to the remote operator that a new message was coming or that transmission of a message was terminated).

Early British Creed machines used the Murray system.

Western Union[edit]

Keyboard of a teleprinter using the Baudot code (US variant), with FIGS and LTRS shift keys

Murray's code was adopted by Western Union which used it until the 1950s, with a few changes that consisted of omitting some characters and adding more control codes. An explicit SPC (space) character was introduced, in place of the BLANK/NULL, and a new BEL code rang a bell or otherwise produced an audible signal at the receiver. Additionally, the WRU or "Who aRe yoU?" code was introduced, which caused a receiving machine to send an identification stream back to the sender.

ITA2[edit]

In 1924 the CCITT introduced the International Telegraph Alphabet No. 2 (ITA2) code[15] as an international standard, which was based on the Western Union code with some minor changes. The US standardized on a version of ITA2 called the American Teletypewriter code (US TTY) which was the basis for 5-bit teletypewriter codes until the debut of 7-bit ASCII in 1963.[16]

International Telegraph Alphabet 2 (British variant)
International telegraphy alphabet No. 2 (Baudot-Murray code)[17]
Impulse patterns
(1=mark, 0=space)
Letter shift Figure shift
Lsb on
right.
Code elements:
543·21
Lsb on
left.
Code elements:
12·345
Punched
marks
ITA2
standard
Russian
MTK-2
variant
Russian
MTK-2
variant
ITA2
standard
US TTY
variant
000·00 00·000 0 Null
010·00 00·010 1 Carriage return
000·10 01·000 1 Line feed
001·00 00·100 1 Space
101·11 11·101 4 Q Я 1
100·11 11·001 3 W В 2
000·01 10·000 1 E Е 3
010·10 01·010 2 R Р 4
100·00 00·001 1 T Т 5
101·01 10·101 3 Y Ы 6
001·11 11·100 3 U У 7
001·10 01·100 2 I И 8
110·00 00·011 2 O О 9
101·10 01·101 3 P П 0
000·11 11·000 2 A А
001·01 10·100 2 S С ' Bell
010·01 10·010 2 D Д WRU? $
011·01 10·110 3 F Ф Э  !
110·10 01·011 3 G Г Ш &
101·00 00·101 2 H Х Щ £ #
010·11 11·010 3 J Й Ю Bell '
011·11 11·110 4 K К (
100·10 01·001 2 L Л )
100·01 10·001 2 Z З + "
111·01 10·111 4 X Ь /
011·10 01·110 3 C Ц  :
111·10 01·111 4 V Ж =  ;
110·01 10·011 3 B Б  ?
011·00 00·110 2 N Н ,
111·00 00·111 3 M М .
110·11 11·011 4 Shift to Figures (FS) Reserved for
figures extension
111·11 11·111 5 Reserved for
lettercase extension
Shift to Letters (LS)
/ Erasure / Delete

The code position assigned to Null was in fact used only for the idle state of teleprinters. During long periods of idle time, the impulse rate was not synchronized between both devices (which could even be powered off or not permanently interconnected on commuted phone lines). To start a message it was first necessary to calibrate the impulse rate a sequence of regularly timed "mark" pulses (1) by group of 5 pulses, which could also be detected by simple passive electronic devices to turn on the teleprinter; this series of pulse was generating series of Erasure/Delete and also initializing the receiver state to the Letters shift mode, however the first pulse could be lost, so this power on procedure could then be terminated by a single Null immediately followed by an Erasure/Delete character. To preserve the synchronization between devices, the Null code could not be used arbitrarily in the middle of messages (this was an improvement to the initial Baudot system where spaces were not explicitly differentiated, so it was difficult to maintain the pulse counters for repeating spaces on teleprinters). But it was then possible to resynchronize devices at any time by sending a Null in the middle of a message (immediately followed by an Erasure/Delete/LS control if followed by a letter, or by a FS control if followed by a figure). Sending Null controls also did not cause the paper band to advance to the next row (as nothing was punched), so this saved precious lengths of punchable paper band. On the opposite the Erasure/Delete/LS control code was always punched and always shifted to the (initial) letters mode.

The Shift to Letters code (LS) is also usable as a way to cancel/delete text from a punched tape after it has been read, allowing a safe destruction of the message before recycling the punched band. For that function, it also plays the same role of filler as the Delete code in ASCII (and in other 7-bit or 8-bit encodings, including EBCDIC for punched cards). Once codes for a fragment text has been replaced by arbitrary number of LS codes, what follows is still preserved and decodable. It can also be used as an initiator to make sure that the decoding of the first code will not give a digit or another symbol from the figures page (because the Null code may be arbitrarily inserted near the end of a punch band or at start of it, and has to be ignored, whereas the Space code is significant in text).

The cells marked as reserved for extensions (using the LS code again from the letters shift page, just after a first LS code to shift from the figures page) has been defined to shift into a new mode: in this new mode, the letters page are containing lowercase letters only, but a third page of codes is still accessible for the uppercase letters, either temporarily for a single letter (encode LS before that letter), may be locked (with FS+LS) for an unlimited number of capital letters or digits and then unlocked to return to lowercase mode (with a single LS).[18] The cell marked as "Reserved" is also usable (using the FS code from the figures shift page) to switch the page of figures (which normally contains digits and national lowercase letters or symbols) to a fourth page (where national letters are uppercased and other symbols may be encoded).

ITA2 is still used in Telecommunications devices for the deaf (TDD), telex, and some amateur radio applications, such as radioteletype ("RTTY"). ITA2 is also used in Enhanced Broadcast Solution (an early 21st century financial protocol specified by Deutsche Börse) to reduce the character encoding footprint.[19]

Nomenclature[edit]

Nearly all 20th-century teleprinter equipment used Western Union's code, ITA2, or variants thereof. Radio amateurs casually call ITA2 and variants "Baudot" incorrectly,[20] and even the American Radio Relay League's Amateur Radio Handbook does so, though in more recent editions the tables of codes correctly identifies it as ITA2.

Details[edit]

NOTE: This table presumes the space called "1" by Baudot and Murray is rightmost, and least significant. The way the transmitted bits were packed into larger codes varied by manufacturer; the most common solution allocates the bits from the least significant bit towards the most significant bit (leaving the three most significant bits of a byte unused).

Table of ITA2 codes (expressed as hexadecimal numbers)

In ITA2, characters are expressed using five bits. ITA2 uses two code sub-sets, the "letter shift" (LTRS), and the "figure shift" (FIGS). The FIGS character (11011) signals that the following characters are to be interpreted as being in the FIGS set, until this is reset by the LTRS (11111) character. In use, the LTRS or FIGS shift key is pressed and released, transmitting the corresponding shift character to the other machine. The desired letters or figures characters are then typed. Unlike a typewriter or modern computer keyboard the shift key isn't kept depressed whilst the corresponding characters are typed. "ENQuiry" will trigger the other machine's answerback. It means "Who are you?"

CR is carriage return, LF is line feed, BEL is the bell character which rang a small bell (often used to alert operators to an incoming message), SP is space, and NUL is the null character (blank tape).

Note: the binary conversions of the codepoints are often shown in reverse order, depending on (presumably) from which side one views the paper tape. Note further that the "control" characters were chosen so that they were either symmetric or in useful pairs so that inserting a tape "upside down" did not result in problems for the equipment and the resulting printout could be deciphered. Thus FIGS (11011), LTRS (11111) and space (00100) are invariant, while CR (00010) and LF (01000), generally used as a pair, are treated the same regardless of order by page printers.[21] LTRS could also be used to overpunch characters to be deleted on a paper tape (much like DEL in 7-bit ASCII).

The sequence RYRYRY... is often used in test messages, and at the start of every transmission. Since R is 01010 and Y is 10101, the sequence exercises much of a teleprinter's mechanical components at maximum stress. Also, at one time, fine-tuning of the receiver was done using two coloured lights (one for each tone). 'RYRYRY...' produced 0101010101..., which made the lights glow with equal brightness when the tuning was correct. This tuning sequence is only useful when ITA2 is used with two-tone FSK modulation, such as is commonly seen in Radioteletype (RTTY) usage.

US implementations of Baudot code may differ in the addition of a few characters, such as #, & on the FIGS layer.

The Russian version of Baudot code (MTK-2) used three shift modes; the Cyrillic letter mode was activated by the character (00000). Because of the larger number of characters in the Cyrillic alphabet, the characters !, &, £ were omitted and replaced by Cyrillics, and BEL has the same code as Cyrillic letter Ю.

See also[edit]

Notes[edit]

  1. ^ Ralston, Anthony; Reilly, Edwin D., eds. (1993), "Baudot Code", Encyclopedia of Computer Science (Third ed.), New York: IEEE Press/Van Nostrand Reinhold, ISBN 0-442-27679-6 
  2. ^ Bacon's Bilateral Cipher (PDF), retrieved 15 April 2012 
  3. ^ H. A. Emmons (1 May 1916). "Printer Systems". Wire & Radio Communications. 34: 209. 
  4. ^ William V. Vansize (25 Jan 1901). "A New Page-Printing Telegraph". Transactions. American Institute of Electrical Engineers. 18: 22. 
  5. ^ "Gauss-Weber-Telegraph". Metrology Mile (in German). Measurement Valley. Retrieved 2009-05-03. 
  6. ^ Pickover, Clifford A. (2009). The Math Book: From Pythagoras to the 57th Dimension, 250 Milestones in the History of Mathematics. Sterling Publishing Company. p. 392. 
  7. ^ Procès d'Amiens Baudot vs Mimault
  8. ^ a b Jennings 2004
  9. ^ Beauchamp, K.G. (2001). History of Telegraphy: Its Technology and Application. Institution of Engineering and Technology. pp. 394–395. ISBN 0-85296-792-6. 
  10. ^ Alan G. Hobbs, 5 Unit Codes, section Baudot Multiplex System
  11. ^ Gleick, James (2011). The Information: A History, a Theory, a Flood. London: Fourth Estate. p. 203. ISBN 978-0-00-742311-8. 
  12. ^ Foster, Maximilian (August 1901). "A Successful Printing Telegraph". The World's Work: A History of Our Time. II: 1195–1199. Retrieved 2009-07-09. 
  13. ^ Copeland 2006, p. 38
  14. ^ Telegraph and Telephone Age. 1921. Donald Murray wrote: 'I allocated the most frequently used letters in English language to the signals represented by the fewest holes in the perforated tape, and so on in proportion.' 
  15. ^ "BruXy: Radio Teletype communication". 2005-10-10. Retrieved 2016-05-09. The transmitted code use International Telegraph Alphabet No. 2 (ITA-2) which was introduced by CCITT in 1924. 
  16. ^ Smith, Gil (2001). "Teletype Communication Codes" (PDF). Baudot.net. Archived (PDF) from the original on 20 August 2008. Retrieved 2008-07-11. 
  17. ^ dataIP Limited. "The "Baudot" Code". Archived from the original on 26 August 2010. Retrieved 9 October 2010 
  18. ^ ITU-T Recommendation S.2 / 11/1988, published in Fascicle VII.1 of the Blue Book
  19. ^ "Enhanced Broadcast Solution – Interface Specification Final Version" (PDF). Deutsche Börse. 17 May 2010. Retrieved 10 August 2011. 
  20. ^ Gillam, Richard (2002). Unicode Demystified:. Addison-Wesley. p. 30. ISBN 0-201-70052-2.  Enhanced Broadcast Solution – Interface Specification Final Version
  21. ^ Jennings, Tom (29 October 2004). "An annotated history of some character codes". Retrieved 22 August 2013. 

References[edit]

This article is based on material taken from the Free On-line Dictionary of Computing prior to 1 November 2008 and incorporated under the "relicensing" terms of the GFDL, version 1.3 or later.