Encode::JP β Japanese Encodings
| Use Case | Command | Description |
|---|---|---|
| Encode UTF-8 to EUC-JP | encode("euc-jp", $utf8) | π Converts UTF-8 string to EUC-JP encoding |
| Decode EUC-JP to UTF-8 | decode("euc-jp", $euc_jp) | π Converts EUC-JP string back to UTF-8 |
| Encode UTF-8 to Shift JIS | encode("shiftjis", $utf8) | π Converts UTF-8 to Shift JIS (MS Kanji) |
| Decode Shift JIS to UTF-8 | decode("shiftjis", $sjis) | π Converts Shift JIS back to UTF-8 |
| Encode UTF-8 to ISO-2022-JP | encode("iso-2022-jp", $utf8) | π§ Converts UTF-8 to 7βbit JIS for email |
| Decode ISO-2022-JP-1 | decode("iso-2022-jp-1", $stream) | π§ Handles extended JIS X 0212 characters |
use Encode qw/encode decode/;
$euc_jp = encode("euc-jp", $utf8); # loads Encode::JP implicitly
$utf8 = decode("euc-jp", $euc_jp); # ditto
This module implements Japanese charset encodings. Encodings supported are as follows.
| Canonical | Alias | Description |
|---|---|---|
| euc-jp | /\beuc.*jp$/i/\bjp.*euc/i/\bujis$/i | EUC (Extended Unix Character) π |
| shiftjis | /\bshift.*jis$/i/\bsjis$/i | Shift JIS (aka MS Kanji) π |
| 7bit-jis | /\bjis$/i | 7βbit JIS |
| iso-2022-jp | ISO-2022-JP [RFC1468] π§ = 7βbit JIS with all Halfwidth Kana converted to Fullwidth | |
| iso-2022-jp-1 | ISO-2022-JP-1 [RFC2237] π§ = ISO-2022-JP with JIS X 0212-1990 support. See below | |
| MacJapanese | Shift JIS + Apple vendor mappings π | |
| cp932 | /\bwindows-31j$/i | Code Page 932 πͺ = Shift JIS + MS/IBM vendor mappings |
| jis0201-raw | JIS0201, raw format | |
| jis0208-raw | JIS0208, raw format | |
| jis0212-raw | JIS0212, raw format |
To find out how to use this module in detail, see Encode.
ISO-2022-JP-1 (RFC2237) is a superset of ISO-2022-JP (RFC1468) which adds support for JIS X 0212-1990. That means you can use the same code to decode to utf8 but not vice versa.
$utf8 = decode('iso-2022-jp-1', $stream);
and
$utf8 = decode('iso-2022-jp', $stream);
yield the same result but
$with_0212 = encode('iso-2022-jp-1', $utf8);
is now different from
$without_0212 = encode('iso-2022-jp', $utf8 );
In the latter case, characters that map to 0212 are first converted to U+3013 (0xA2AE in EUC-JP; a white square also known as βTofuβ or βgeta markβ) then fed to the decoding engine. U+FFFD is not used, in order to preserve text layout as much as possible.
The ASCII region (0x00-0x7f) is preserved for all encodings, even though this conflicts with mappings by the Unicode Consortium.
Generated by phpman v4.9.26-5-g7740029 · Markdown · JSON · MCP Author: Che Dong Under GNU General Public License
2026-08-11 04:24 @216.73.216.11
CrawledBy Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)