perldoc > Encode::JP

πŸ“– NAME

Encode::JP β€” Japanese Encodings

πŸš€ Quick Reference

Use CaseCommandDescription
Encode UTF-8 to EUC-JPencode("euc-jp", $utf8)🌐 Converts UTF-8 string to EUC-JP encoding
Decode EUC-JP to UTF-8decode("euc-jp", $euc_jp)🌐 Converts EUC-JP string back to UTF-8
Encode UTF-8 to Shift JISencode("shiftjis", $utf8)πŸ“„ Converts UTF-8 to Shift JIS (MS Kanji)
Decode Shift JIS to UTF-8decode("shiftjis", $sjis)πŸ“„ Converts Shift JIS back to UTF-8
Encode UTF-8 to ISO-2022-JPencode("iso-2022-jp", $utf8)πŸ“§ Converts UTF-8 to 7‑bit JIS for email
Decode ISO-2022-JP-1decode("iso-2022-jp-1", $stream)πŸ“§ Handles extended JIS X 0212 characters

πŸ› οΈ SYNOPSIS

use Encode qw/encode decode/;
$euc_jp = encode("euc-jp", $utf8);   # loads Encode::JP implicitly
$utf8   = decode("euc-jp", $euc_jp); # ditto

πŸ“ ABSTRACT

This module implements Japanese charset encodings. Encodings supported are as follows.

CanonicalAliasDescription
euc-jp/\beuc.*jp$/i
/\bjp.*euc/i
/\bujis$/i
EUC (Extended Unix Character) 🌐
shiftjis/\bshift.*jis$/i
/\bsjis$/i
Shift JIS (aka MS Kanji) πŸ“„
7bit-jis/\bjis$/i7‑bit JIS
iso-2022-jpISO-2022-JP [RFC1468] πŸ“§
= 7‑bit JIS with all Halfwidth Kana converted to Fullwidth
iso-2022-jp-1ISO-2022-JP-1 [RFC2237] πŸ“§
= ISO-2022-JP with JIS X 0212-1990 support. See below
MacJapaneseShift JIS + Apple vendor mappings 🍏
cp932/\bwindows-31j$/iCode Page 932 πŸͺŸ
= Shift JIS + MS/IBM vendor mappings
jis0201-rawJIS0201, raw format
jis0208-rawJIS0208, raw format
jis0212-rawJIS0212, raw format

πŸ“š DESCRIPTION

To find out how to use this module in detail, see Encode.

πŸ“Œ Note on ISO-2022-JP(-1)

ISO-2022-JP-1 (RFC2237) is a superset of ISO-2022-JP (RFC1468) which adds support for JIS X 0212-1990. That means you can use the same code to decode to utf8 but not vice versa.

$utf8 = decode('iso-2022-jp-1', $stream);

and

$utf8 = decode('iso-2022-jp',   $stream);

yield the same result but

$with_0212 = encode('iso-2022-jp-1', $utf8);

is now different from

$without_0212 = encode('iso-2022-jp', $utf8 );

In the latter case, characters that map to 0212 are first converted to U+3013 (0xA2AE in EUC-JP; a white square also known as β€˜Tofu’ or β€˜geta mark’) then fed to the decoding engine. U+FFFD is not used, in order to preserve text layout as much as possible.

πŸ› BUGS

The ASCII region (0x00-0x7f) is preserved for all encodings, even though this conflicts with mappings by the Unicode Consortium.

πŸ”— SEE ALSO

Encode

Encode::JP
πŸ“– NAME πŸš€ Quick Reference πŸ› οΈ SYNOPSIS πŸ“ ABSTRACT πŸ“š DESCRIPTION
πŸ“Œ Note on ISO-2022-JP(-1)
πŸ› BUGS πŸ”— SEE ALSO

Generated by phpman v4.9.26-5-g7740029 · Markdown · JSON · MCP Author: Che Dong Under GNU General Public License
2026-08-11 04:24 @216.73.216.11
CrawledBy Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)
Valid XHTML 1.0 Transitional!Valid CSS!