{
    "mode": "perldoc",
    "parameter": "Locale::Maketext::TPJ13",
    "section": "",
    "url": "https://www.chedong.com/phpMan.php/perldoc/Locale%3A%3AMaketext%3A%3ATPJ13/json",
    "generated": "2026-10-04T17:51:26Z",
    "synopsis": "# This an article, not a module.",
    "sections": {
        "NAME": {
            "content": "Locale::Maketext::TPJ13 -- article about software localization\n",
            "subsections": []
        },
        "SYNOPSIS": {
            "content": "# This an article, not a module.\n",
            "subsections": []
        },
        "DESCRIPTION": {
            "content": "The following article by Sean M. Burke and Jordan Lachler first appeared in *The Perl Journal*\n#13 and is copyright 1999 The Perl Journal. It appears courtesy of Jon Orwant and The Perl\nJournal. This document may be distributed under the same terms as Perl itself.\n",
            "subsections": []
        },
        "Localization and Perl: gettext breaks, Maketext fixes": {
            "content": "by Sean M. Burke and Jordan Lachler\n\nThis article points out cases where gettext (a common system for localizing software interfaces\n-- i.e., making them work in the user's language of choice) fails because of basic differences\nbetween human languages. This article then describes Maketext, a new system capable of correctly\ntreating these differences.\n\nA Localization Horror Story: It Could Happen To You\n\"There are a number of languages spoken by human beings in this world.\"\n\n-- Harald Tveit Alvestrand, in RFC 1766, \"Tags for the Identification of Languages\"\n\nImagine that your task for the day is to localize a piece of software -- and luckily for you,\nthe only output the program emits is two messages, like this:\n\nI scanned 12 directories.\n\nYour query matched 10 files in 4 directories.\n\nSo how hard could that be? You look at the code that produces the first item, and it reads:\n\nprintf(\"I scanned %g directories.\",\n$directorycount);\n\nYou think about that, and realize that it doesn't even work right for English, as it can produce\nthis output:\n\nI scanned 1 directories.\n\nSo you rewrite it to read:\n\nprintf(\"I scanned %g %s.\",\n$directorycount,\n$directorycount == 1 ?\n\"directory\" : \"directories\",\n);\n\n...which does the Right Thing. (In case you don't recall, \"%g\" is for locale-specific number\ninterpolation, and \"%s\" is for string interpolation.)\n\nBut you still have to localize it for all the languages you're producing this software for, so\nyou pull Locale::gettext off of CPAN so you can access the \"gettext\" C functions you've heard\nare standard for localization tasks.\n\nAnd you write:\n\nprintf(gettext(\"I scanned %g %s.\"),\n$dirscancount,\n$dirscancount == 1 ?\ngettext(\"directory\") : gettext(\"directories\"),\n);\n\nBut you then read in the gettext manual (Drepper, Miller, and Pinard 1995) that this is not a\ngood idea, since how a single word like \"directory\" or \"directories\" is translated may depend on\ncontext -- and this is true, since in a case language like German or Russian, you'd may need\nthese words with a different case ending in the first instance (where the word is the object of\na verb) than in the second instance, which you haven't even gotten to yet (where the word is the\nobject of a preposition, \"in %g directories\") -- assuming these keep the same syntax when\ntranslated into those languages.\n\nSo, on the advice of the gettext manual, you rewrite:\n\nprintf( $dirscancount == 1 ?\ngettext(\"I scanned %g directory.\") :\ngettext(\"I scanned %g directories.\"),\n$dirscancount );\n\nSo, you email your various translators (the boss decides that the languages du jour are Chinese,\nArabic, Russian, and Italian, so you have one translator for each), asking for translations for\n\"I scanned %g directory.\" and \"I scanned %g directories.\". When they reply, you'll put that in\nthe lexicons for gettext to use when it localizes your software, so that when the user is\nrunning under the \"zh\" (Chinese) locale, gettext(\"I scanned %g directory.\") will return the\nappropriate Chinese text, with a \"%g\" in there where printf can then interpolate $dirscan.\n\nYour Chinese translator emails right back -- he says both of these phrases translate to the same\nthing in Chinese, because, in linguistic jargon, Chinese \"doesn't have number as a grammatical\ncategory\" -- whereas English does. That is, English has grammatical rules that refer to\n\"number\", i.e., whether something is grammatically singular or plural; and one of these rules is\nthe one that forces nouns to take a plural suffix (generally \"s\") when in a plural context, as\nthey are when they follow a number other than \"one\" (including, oddly enough, \"zero\"). Chinese\nhas no such rules, and so has just the one phrase where English has two. But, no problem, you\ncan have this one Chinese phrase appear as the translation for the two English phrases in the\n\"zh\" gettext lexicon for your program.\n\nEmboldened by this, you dive into the second phrase that your software needs to output: \"Your\nquery matched 10 files in 4 directories.\". You notice that if you want to treat phrases as\nindivisible, as the gettext manual wisely advises, you need four cases now, instead of two, to\ncover the permutations of singular and plural on the two items, $dircount and $filecount. So\nyou try this:\n\nprintf( $filecount == 1 ?\n( $directorycount == 1 ?\ngettext(\"Your query matched %g file in %g directory.\") :\ngettext(\"Your query matched %g file in %g directories.\") ) :\n( $directorycount == 1 ?\ngettext(\"Your query matched %g files in %g directory.\") :\ngettext(\"Your query matched %g files in %g directories.\") ),\n$filecount, $directorycount,\n);\n\n(The case of \"1 file in 2 [or more] directories\" could, I suppose, occur in the case of\nsymlinking or something of the sort.)\n\nIt occurs to you that this is not the prettiest code you've ever written, but this seems the way\nto go. You mail off to the translators asking for translations for these four cases. The Chinese\nguy replies with the one phrase that these all translate to in Chinese, and that phrase has two\n\"%g\"s in it, as it should -- but there's a problem. He translates it word-for-word back: \"In %g\ndirectories contains %g files match your query.\" The %g slots are in an order reverse to what\nthey are in English. You wonder how you'll get gettext to handle that.\n\nBut you put it aside for the moment, and optimistically hope that the other translators won't\nhave this problem, and that their languages will be better behaved -- i.e., that they will be\njust like English.\n\nBut the Arabic translator is the next to write back. First off, your code for \"I scanned %g\ndirectory.\" or \"I scanned %g directories.\" assumes there's only singular or plural. But, to use\nlinguistic jargon again, Arabic has grammatical number, like English (but unlike Chinese), but\nit's a three-term category: singular, dual, and plural. In other words, the way you say\n\"directory\" depends on whether there's one directory, or *two* of them, or *more than two* of\nthem. Your test of \"($directory == 1)\" no longer does the job. And it means that where English's\ngrammatical category of number necessitates only the two permutations of the first sentence\nbased on \"directory [singular]\" and \"directories [plural]\", Arabic has three -- and, worse, in\nthe second sentence (\"Your query matched %g file in %g directory.\"), where English has four,\nArabic has nine. You sense an unwelcome, exponential trend taking shape.\n\nYour Italian translator emails you back and says that \"I searched 0 directories\" (a possible\nEnglish output of your program) is stilted, and if you think that's fine English, that's your\nproblem, but that *just will not do* in the language of Dante. He insists that where\n$directorycount is 0, your program should produce the Italian text for \"I *didn't* scan *any*\ndirectories.\". And ditto for \"I didn't match any files in any directories\", although he says the\nlast part about \"in any directories\" should probably just be left off.\n\nYou wonder how you'll get gettext to handle this; to accommodate the ways Arabic, Chinese, and\nItalian deal with numbers in just these few very simple phrases, you need to write code that\nwill ask gettext for different queries depending on whether the numerical values in question are\n1, 2, more than 2, or in some cases 0, and you still haven't figured out the problem with the\ndifferent word order in Chinese.\n\nThen your Russian translator calls on the phone, to *personally* tell you the bad news about how\nreally unpleasant your life is about to become:\n\nRussian, like German or Latin, is an inflectional language; that is, nouns and adjectives have\nto take endings that depend on their case (i.e., nominative, accusative, genitive, etc...) --\nwhich is roughly a matter of what role they have in syntax of the sentence -- as well as on the\ngrammatical gender (i.e., masculine, feminine, neuter) and number (i.e., singular or plural) of\nthe noun, as well as on the declension class of the noun. But unlike with most other inflected\nlanguages, putting a number-phrase (like \"ten\" or \"forty-three\", or their Arabic numeral\nequivalents) in front of noun in Russian can change the case and number that noun is, and\ntherefore the endings you have to put on it.\n\nHe elaborates: In \"I scanned %g directories\", you'd *expect* \"directories\" to be in the\naccusative case (since it is the direct object in the sentence) and the plural number, except\nwhere $directorycount is 1, then you'd expect the singular, of course. Just like Latin or\nGerman. *But!* Where $directorycount % 10 is 1 (\"%\" for modulo, remember), assuming $directory\ncount is an integer, and except where $directorycount % 100 is 11, \"directories\" is forced to\nbecome grammatically singular, which means it gets the ending for the accusative singular... You\nbegin to visualize the code it'd take to test for the problem so far, *and still work for\nChinese and Arabic and Italian*, and how many gettext items that'd take, but he keeps going...\nBut where $directorycount % 10 is 2, 3, or 4 (except where $directorycount % 100 is 12, 13, or\n14), the word for \"directories\" is forced to be genitive singular -- which means another\nending... The room begins to spin around you, slowly at first... But with *all other* integer\nvalues, since \"directory\" is an inanimate noun, when preceded by a number and in the nominative\nor accusative cases (as it is here, just your luck!), it does stay plural, but it is forced into\nthe genitive case -- yet another ending... And you never hear him get to the part about how\nyou're going to run into similar (but maybe subtly different) problems with other Slavic\nlanguages like Polish, because the floor comes up to meet you, and you fade into\nunconsciousness.\n\nThe above cautionary tale relates how an attempt at localization can lead from programmer\nconsternation, to program obfuscation, to a need for sedation. But careful evaluation shows that\nyour choice of tools merely needed further consideration.\n",
            "subsections": [
                {
                    "name": "The Linguistic View",
                    "content": "\"It is more complicated than you think.\"\n\n-- The Eighth Networking Truth, from RFC 1925\n\nThe field of Linguistics has expended a great deal of effort over the past century trying to\nfind grammatical patterns which hold across languages; it's been a constant process of people\nmaking generalizations that should apply to all languages, only to find out that, all too often,\nthese generalizations fail -- sometimes failing for just a few languages, sometimes whole\nclasses of languages, and sometimes nearly every language in the world except English. Broad\nstatistical trends are evident in what the \"average language\" is like as far as what its rules\ncan look like, must look like, and cannot look like. But the \"average language\" is just as\nunreal a concept as the \"average person\" -- it runs up against the fact no language (or person)\nis, in fact, average. The wisdom of past experience leads us to believe that any given language\ncan do whatever it wants, in any order, with appeal to any kind of grammatical categories wants\n-- case, number, tense, real or metaphoric characteristics of the things that words refer to,\narbitrary or predictable classifications of words based on what endings or prefixes they can\ntake, degree or means of certainty about the truth of statements expressed, and so on, ad\ninfinitum.\n\nMercifully, most localization tasks are a matter of finding ways to translate whole phrases,\ngenerally sentences, where the context is relatively set, and where the only variation in\ncontent is *usually* in a number being expressed -- as in the example sentences above.\nTranslating specific, fully-formed sentences is, in practice, fairly foolproof -- which is good,\nbecause that's what's in the phrasebooks that so many tourists rely on. Now, a given phrase\n(whether in a phrasebook or in a gettext lexicon) in one language *might* have a greater or\nlesser applicability than that phrase's translation into another language -- for example,\nstrictly speaking, in Arabic, the \"your\" in \"Your query matched...\" would take a different form\ndepending on whether the user is male or female; so the Arabic translation \"your[feminine]\nquery\" is applicable in fewer cases than the corresponding English phrase, which doesn't\ndistinguish the user's gender. (In practice, it's not feasible to have a program know the user's\ngender, so the masculine \"you\" in Arabic is usually used, by default.)\n\nBut in general, such surprises are rare when entire sentences are being translated, especially\nwhen the functional context is restricted to that of a computer interacting with a user either\nto convey a fact or to prompt for a piece of information. So, for purposes of localization,\ntranslation by phrase (generally by sentence) is both the simplest and the least problematic.\n"
                },
                {
                    "name": "Breaking gettext",
                    "content": "\"It Has To Work.\"\n\n-- First Networking Truth, RFC 1925\n\nConsider that sentences in a tourist phrasebook are of two types: ones like \"How do I get to the\nmarketplace?\" that don't have any blanks to fill in, and ones like \"How much do these\ncost?\", where there's one or more blanks to fill in (and these are usually linked to a list of\nwords that you can put in that blank: \"fish\", \"potatoes\", \"tomatoes\", etc.). The ones with no\nblanks are no problem, but the fill-in-the-blank ones may not be really straightforward. If it's\na Swahili phrasebook, for example, the authors probably didn't bother to tell you the\ncomplicated ways that the verb \"cost\" changes its inflectional prefix depending on the noun\nyou're putting in the blank. The trader in the marketplace will still understand what you're\nsaying if you say \"how much do these potatoes cost?\" with the wrong inflectional prefix on\n\"cost\". After all, *you* can't speak proper Swahili, *you're* just a tourist. But while tourists\ncan be stupid, computers are supposed to be smart; the computer should be able to fill in the\nblank, and still have the results be grammatical.\n\nIn other words, a phrasebook entry takes some values as parameters (the things that you fill in\nthe blank or blanks), and provides a value based on these parameters, where the way you get that\nfinal value from the given values can, properly speaking, involve an arbitrarily complex series\nof operations. (In the case of Chinese, it'd be not at all complex, at least in cases like the\nexamples at the beginning of this article; whereas in the case of Russian it'd be a rather\ncomplex series of operations. And in some languages, the complexity could be spread around\ndifferently: while the act of putting a number-expression in front of a noun phrase might not be\ncomplex by itself, it may change how you have to, for example, inflect a verb elsewhere in the\nsentence. This is what in syntax is called \"long-distance dependencies\".)\n\nThis talk of parameters and arbitrary complexity is just another way to say that an entry in a\nphrasebook is what in a programming language would be called a \"function\". Just so you don't\nmiss it, this is the crux of this article: *A phrase is a function; a phrasebook is a bunch of\nfunctions.*\n\nThe reason that using gettext runs into walls (as in the above second-person horror story) is\nthat you're trying to use a string (or worse, a choice among a bunch of strings) to do what you\nreally need a function for -- which is futile. Preforming (s)printf interpolation on the strings\nwhich you get back from gettext does allow you to do *some* common things passably well...\nsometimes... sort of; but, to paraphrase what some people say about \"csh\" script programming,\n\"it fools you into thinking you can use it for real things, but you can't, and you don't\ndiscover this until you've already spent too much time trying, and by then it's too late.\"\n"
                },
                {
                    "name": "Replacing gettext",
                    "content": "So, what needs to replace gettext is a system that supports lexicons of functions instead of\nlexicons of strings. An entry in a lexicon from such a system should *not* look like this:\n\n\"J'ai trouv\\xE9 %g fichiers dans %g r\\xE9pertoires\"\n\n[\\xE9 is e-acute in Latin-1. Some pod renderers would scream if I used the actual character\nhere. -- SB]\n\nbut instead like this, bearing in mind that this is just a first stab:\n\nsub IfoundX1filesinX2directories {\nmy( $files, $dirs ) = @[0,1];\n$files = sprintf(\"%g %s\", $files,\n$files == 1 ? 'fichier' : 'fichiers');\n$dirs = sprintf(\"%g %s\", $dirs,\n$dirs == 1 ? \"r\\xE9pertoire\" : \"r\\xE9pertoires\");\nreturn \"J'ai trouv\\xE9 $files dans $dirs.\";\n}\n\nNow, there's no particularly obvious way to store anything but strings in a gettext lexicon; so\nit looks like we just have to start over and make something better, from scratch. I call my shot\nat a gettext-replacement system \"Maketext\", or, in CPAN terms, Locale::Maketext.\n\nWhen designing Maketext, I chose to plan its main features in terms of \"buzzword compliance\".\nAnd here are the buzzwords:\n"
                },
                {
                    "name": "Buzzwords: Abstraction and Encapsulation",
                    "content": "The complexity of the language you're trying to output a phrase in is entirely abstracted inside\n(and encapsulated within) the Maketext module for that interface. When you call:\n\nprint $lang->maketext(\"You have [quant,1,piece] of new mail.\",\nscalar(@messages));\n\nyou don't know (and in fact can't easily find out) whether this will involve lots of figuring,\nas in Russian (if $lang is a handle to the Russian module), or relatively little, as in Chinese.\nThat kind of abstraction and encapsulation may encourage other pleasant buzzwords like\nmodularization and stratification, depending on what design decisions you make.\n"
                },
                {
                    "name": "Buzzword: Isomorphism",
                    "content": "\"Isomorphism\" means \"having the same structure or form\"; in discussions of program design, the\nword takes on the special, specific meaning that your implementation of a solution to a problem\n*has the same structure* as, say, an informal verbal description of the solution, or maybe of\nthe problem itself. Isomorphism is, all things considered, a good thing -- it's what\nproblem-solving (and solution-implementing) should look like.\n\nWhat's wrong the with gettext-using code like this...\n\nprintf( $filecount == 1 ?\n( $directorycount == 1 ?\n\"Your query matched %g file in %g directory.\" :\n\"Your query matched %g file in %g directories.\" ) :\n( $directorycount == 1 ?\n\"Your query matched %g files in %g directory.\" :\n\"Your query matched %g files in %g directories.\" ),\n$filecount, $directorycount,\n);\n\nis first off that it's not well abstracted -- these ways of testing for grammatical number (as\nin the expressions like \"foo == 1 ? singularform : pluralform\") should be abstracted to each\nlanguage module, since how you get grammatical number is language-specific.\n\nBut second off, it's not isomorphic -- the \"solution\" (i.e., the phrasebook entries) for Chinese\nmaps from these four English phrases to the one Chinese phrase that fits for all of them. In\nother words, the informal solution would be \"The way to say what you want in Chinese is with the\none phrase 'For your question, in Y directories you would find X files'\" -- and so the\nimplemented solution should be, isomorphically, just a straightforward way to spit out that one\nphrase, with numerals properly interpolated. It shouldn't have to map from the complexity of\nother languages to the simplicity of this one.\n"
                },
                {
                    "name": "Buzzword: Inheritance",
                    "content": "There's a great deal of reuse possible for sharing of phrases between modules for related\ndialects, or for sharing of auxiliary functions between related languages. (By \"auxiliary\nfunctions\", I mean functions that don't produce phrase-text, but which, say, return an answer to\n\"does this number require a plural noun after it?\". Such auxiliary functions would be used in\nthe internal logic of functions that actually do produce phrase-text.)\n\nIn the case of sharing phrases, consider that you have an interface already localized for\nAmerican English (probably by having been written with that as the native locale, but that's\nincidental). Localizing it for UK English should, in practical terms, be just a matter of\nrunning it past a British person with the instructions to indicate what few phrases would\nbenefit from a change in spelling or possibly minor rewording. In that case, you should be able\nto put in the UK English localization module *only* those phrases that are UK-specific, and for\nall the rest, *inherit* from the American English module. (And I expect this same situation\nwould apply with Brazilian and Continental Portugese, possibly with some *very* closely related\nlanguages like Czech and Slovak, and possibly with the slightly different \"versions\" of written\nMandarin Chinese, as I hear exist in Taiwan and mainland China.)\n\nAs to sharing of auxiliary functions, consider the problem of Russian numbers from the beginning\nof this article; obviously, you'd want to write only once the hairy code that, given a numeric\nvalue, would return some specification of which case and number a given quantified noun should\nuse. But suppose that you discover, while localizing an interface for, say, Ukrainian (a Slavic\nlanguage related to Russian, spoken by several million people, many of whom would be relieved to\nfind that your Web site's or software's interface is available in their language), that the\nrules in Ukrainian are the same as in Russian for quantification, and probably for many other\ngrammatical functions. While there may well be no phrases in common between Russian and\nUkrainian, you could still choose to have the Ukrainian module inherit from the Russian module,\njust for the sake of inheriting all the various grammatical methods. Or, probably better\norganizationally, you could move those functions to a module called \"ESlavic\" or something,\nwhich Russian and Ukrainian could inherit useful functions from, but which would (presumably)\nprovide no lexicon.\n"
                },
                {
                    "name": "Buzzword: Concision",
                    "content": "Okay, concision isn't a buzzword. But it should be, so I decree that as a new buzzword,\n\"concision\" means that simple common things should be expressible in very few lines (or maybe\neven just a few characters) of code -- call it a special case of \"making simple things easy and\nhard things possible\", and see also the role it played in the MIDI::Simple language, discussed\nelsewhere in this issue [TPJ#13].\n\nConsider our first stab at an entry in our \"phrasebook of functions\":\n\nsub IfoundX1filesinX2directories {\nmy( $files, $dirs ) = @[0,1];\n$files = sprintf(\"%g %s\", $files,\n$files == 1 ? 'fichier' : 'fichiers');\n$dirs = sprintf(\"%g %s\", $dirs,\n$dirs == 1 ? \"r\\xE9pertoire\" : \"r\\xE9pertoires\");\nreturn \"J'ai trouv\\xE9 $files dans $dirs.\";\n}\n\nYou may sense that a lexicon (to use a non-committal catch-all term for a collection of things\nyou know how to say, regardless of whether they're phrases or words) consisting of functions\n*expressed* as above would make for rather long-winded and repetitive code -- even if you wisely\nrewrote this to have quantification (as we call adding a number expression to a noun phrase) be\na function called like:\n\nsub IfoundX1filesinX2directories {\nmy( $files, $dirs ) = @[0,1];\n$files = quant($files, \"fichier\");\n$dirs =  quant($dirs,  \"r\\xE9pertoire\");\nreturn \"J'ai trouv\\xE9 $files dans $dirs.\";\n}\n\nAnd you may also sense that you do not want to bother your translators with having to write Perl\ncode -- you'd much rather that they spend their *very costly time* on just translation. And this\nis to say nothing of the near impossibility of finding a commercial translator who would know\neven simple Perl.\n\nIn a first-hack implementation of Maketext, each language-module's lexicon looked like this:\n\n%Lexicon = (\n\"I found %g files in %g directories\"\n=> sub {\nmy( $files, $dirs ) = @[0,1];\n$files = quant($files, \"fichier\");\n$dirs =  quant($dirs,  \"r\\xE9pertoire\");\nreturn \"J'ai trouv\\xE9 $files dans $dirs.\";\n},\n... and so on with other phrase => sub mappings ...\n);\n\nbut I immediately went looking for some more concise way to basically denote the same\nphrase-function -- a way that would also serve to concisely denote *most* phrase-functions in\nthe lexicon for *most* languages. After much time and even some actual thought, I decided on\nthis system:\n\n* Where a value in a %Lexicon hash is a contentful string instead of an anonymous sub (or,\nconceivably, a coderef), it would be interpreted as a sort of shorthand expression of what the\nsub does. When accessed for the first time in a session, it is parsed, turned into Perl code,\nand then eval'd into an anonymous sub; then that sub replaces the original string in that\nlexicon. (That way, the work of parsing and evaling the shorthand form for a given phrase is\ndone no more than once per session.)\n\n* Calls to \"maketext\" (as Maketext's main function is called) happen thru a \"language session\nhandle\", notionally very much like an IO handle, in that you open one at the start of the\nsession, and use it for \"sending signals\" to an object in order to have it return the text you\nwant.\n\nSo, this:\n\n$lang->maketext(\"You have [quant,1,piece] of new mail.\",\nscalar(@messages));\n\nbasically means this: look in the lexicon for $lang (which may inherit from any number of other\nlexicons), and find the function that we happen to associate with the string \"You have\n[quant,1,piece] of new mail\" (which is, and should be, a functioning \"shorthand\" for this\nfunction in the native locale -- English in this case). If you find such a function, call it\nwith $lang as its first parameter (as if it were a method), and then a copy of scalar(@messages)\nas its second, and then return that value. If that function was found, but was in string\nshorthand instead of being a fully specified function, parse it and make it into a function\nbefore calling it the first time.\n\n* The shorthand uses code in brackets to indicate method calls that should be performed. A full\nexplanation is not in order here, but a few examples will suffice:\n\n\"You have [quant,1,piece] of new mail.\"\n\nThe above code is shorthand for, and will be interpreted as, this:\n\nsub {\nmy $handle = $[0];\nmy(@params) = @;\nreturn join '',\n\"You have \",\n$handle->quant($params[1], 'piece'),\n\"of new mail.\";\n}\n\nwhere \"quant\" is the name of a method you're using to quantify the noun \"piece\" with the number\n$params[0].\n\nA string with no brackety calls, like this:\n\n\"Your search expression was malformed.\"\n\nis somewhat of a degenerate case, and just gets turned into:\n\nsub { return \"Your search expression was malformed.\" }\n\nHowever, not everything you can write in Perl code can be written in the above shorthand system\n-- not by a long shot. For example, consider the Italian translator from the beginning of this\narticle, who wanted the Italian for \"I didn't find any files\" as a special case, instead of \"I\nfound 0 files\". That couldn't be specified (at least not easily or simply) in our shorthand\nsystem, and it would have to be written out in full, like this:\n\nsub {  # pretend the English strings are in Italian\nmy($handle, $files, $dirs) = @[0,1,2];\nreturn \"I didn't find any files\" unless $files;\nreturn join '',\n\"I found \",\n$handle->quant($files, 'file'),\n\" in \",\n$handle->quant($dirs,  'directory'),\n\".\";\n}\n\nNext to a lexicon full of shorthand code, that sort of sticks out like a sore thumb -- but this\n*is* a special case, after all; and at least it's possible, if not as concise as usual.\n\nAs to how you'd implement the Russian example from the beginning of the article, well, There's\nMore Than One Way To Do It, but it could be something like this (using English words for\nRussian, just so you know what's going on):\n\n\"I [quant,1,directory,accusative] scanned.\"\n\nThis shifts the burden of complexity off to the quant method. That method's parameters are: the\nnumeric value it's going to use to quantify something; the Russian word it's going to quantify;\nand the parameter \"accusative\", which you're using to mean that this sentence's syntax wants a\nnoun in the accusative case there, although that quantification method may have to overrule, for\ngrammatical reasons you may recall from the beginning of this article.\n\nNow, the Russian quant method here is responsible not only for implementing the strange logic\nnecessary for figuring out how Russian number-phrases impose case and number on their\nnoun-phrases, but also for inflecting the Russian word for \"directory\". How that inflection is\nto be carried out is no small issue, and among the solutions I've seen, some (like variations on\na simple lookup in a hash where all possible forms are provided for all necessary words) are\nstraightforward but *can* become cumbersome when you need to inflect more than a few dozen\nwords; and other solutions (like using algorithms to model the inflections, storing only root\nforms and irregularities) *can* involve more overhead than is justifiable for all but the\nlargest lexicons.\n\nMercifully, this design decision becomes crucial only in the hairiest of inflected languages, of\nwhich Russian is by no means the *worst* case scenario, but is worse than most. Most languages\nhave simpler inflection systems; for example, in English or Swahili, there are generally no more\nthan two possible inflected forms for a given noun (\"error/errors\"; \"kosa/makosa\"), and the\nrules for producing these forms are fairly simple -- or at least, simple rules can be formulated\nthat work for most words, and you can then treat the exceptions as just \"irregular\", at least\nrelative to your ad hoc rules. A simpler inflection system (simpler rules, fewer forms) means\nthat design decisions are less crucial to maintaining sanity, whereas the same decisions could\nincur overhead-versus-scalability problems in languages like Russian. It may *also* be likely\nthat code (possibly in Perl, as with Lingua::EN::Inflect, for English nouns) has already been\nwritten for the language in question, whether simple or complex.\n\nMoreover, a third possibility may even be simpler than anything discussed above: \"Just require\nthat all possible (or at least applicable) forms be provided in the call to the given language's\nquant method, as in:\"\n\n\"I found [quant,1,file,files].\"\n\nThat way, quant just has to chose which form it needs, without having to look up or generate\nanything. While possibly not optimal for Russian, this should work well for most other\nlanguages, where quantification is not as complicated an operation.\n"
                },
                {
                    "name": "The Devil in the Details",
                    "content": "There's plenty more to Maketext than described above -- for example, there's the details of how\nlanguage tags (\"en-US\", \"i-pwn\", \"fi\", etc.) or locale IDs (\"enUS\") interact with actual module\nnaming (\"BogoQuery/Locale/enus.pm\"), and what magic can ensue; there's the details of how to\nrecord (and possibly negotiate) what character encoding Maketext will return text in (UTF8?\nLatin-1? KOI8?). There's the interesting fact that Maketext is for localization, but nowhere\nactually has a \"\"use locale;\"\" anywhere in it. For the curious, there's the somewhat frightening\ndetails of how I actually implement something like data inheritance so that searches across\nmodules' %Lexicon hashes can parallel how Perl implements method inheritance.\n\nAnd, most importantly, there's all the practical details of how to actually go about deriving\nfrom Maketext so you can use it for your interfaces, and the various tools and conventions for\nstarting out and maintaining individual language modules.\n\nThat is all covered in the documentation for Locale::Maketext and the modules that come with it,\navailable in CPAN. After having read this article, which covers the why's of Maketext, the\ndocumentation, which covers the how's of it, should be quite straightforward.\n"
                },
                {
                    "name": "The Proof in the Pudding: Localizing Web Sites",
                    "content": "Maketext and gettext have a notable difference: gettext is in C, accessible thru C library\ncalls, whereas Maketext is in Perl, and really can't work without a Perl interpreter (although I\nsuppose something like it could be written for C). Accidents of history (and not necessarily\nlucky ones) have made C++ the most common language for the implementation of applications like\nword processors, Web browsers, and even many in-house applications like custom query systems.\nCurrent conditions make it somewhat unlikely that the next one of any of these kinds of\napplications will be written in Perl, albeit clearly more for reasons of custom and inertia than\nout of consideration of what is the right tool for the job.\n\nHowever, other accidents of history have made Perl a well-accepted language for design of\nserver-side programs (generally in CGI form) for Web site interfaces. Localization of static\npages in Web sites is trivial, feasible either with simple language-negotiation features in\nservers like Apache, or with some kind of server-side inclusions of language-appropriate text\ninto layout templates. However, I think that the localization of Perl-based search systems (or\nother kinds of dynamic content) in Web sites, be they public or access-restricted, is where\nMaketext will see the greatest use.\n\nI presume that it would be only the exceptional Web site that gets localized for English *and*\nChinese *and* Italian *and* Arabic *and* Russian, to recall the languages from the beginning of\nthis article -- to say nothing of German, Spanish, French, Japanese, Finnish, and Hindi, to name\na few languages that benefit from large numbers of programmers or Web viewers or both.\n\nHowever, the ever-increasing internationalization of the Web (whether measured in terms of\namount of content, of numbers of content writers or programmers, or of size of content\naudiences) makes it increasingly likely that the interface to the average Web-based dynamic\ncontent service will be localized for two or maybe three languages. It is my hope that Maketext\nwill make that task as simple as possible, and will remove previous barriers to localization for\nlanguages dissimilar to English.\n\nEND\n\nSean M. Burke (sburke@cpan.org) has a Master's in linguistics from Northwestern University; he\nspecializes in language technology. Jordan Lachler (lachler@unm.edu) is a PhD student in the\nDepartment of Linguistics at the University of New Mexico; he specializes in morphology and\npedagogy of North American native languages.\n"
                },
                {
                    "name": "References",
                    "content": "Alvestrand, Harald Tveit. 1995. *RFC 1766: Tags for the Identification of Languages.*\n\"<http://www.ietf.org/rfc/rfc1766.txt>\" [Now see RFC 3066.]\n\nCallon, Ross, editor. 1996. *RFC 1925: The Twelve Networking Truths.*\n\"<http://www.ietf.org/rfc/rfc1925.txt>\"\n\nDrepper, Ulrich, Peter Miller, and François Pinard. 1995-2001. GNU \"gettext\". Available in\n\"<ftp://prep.ai.mit.edu/pub/gnu/>\", with extensive docs in the distribution tarball. [Since I\nwrote this article in 1998, I now see that the gettext docs are now trying more to come to terms\nwith plurality. Whether useful conclusions have come from it is another question altogether. --\nSMB, May 2001]\n\nForbes, Nevill. 1964. *Russian Grammar.* Third Edition, revised by J. C. Dumbreck. Oxford\nUniversity Press.\n"
                }
            ]
        }
    },
    "summary": "Locale::Maketext::TPJ13 -- article about software localization",
    "flags": [],
    "examples": [],
    "see_also": []
}