Sample code for 30+ languages & platforms
AutoIt Requires Chilkat v11.1.0+

Unicode Escape and Unescape Text in StringBuilder

Demonstrates options for unicode escaping non-us-ascii chars and emojis.

Chilkat AutoIt Downloads

AutoIt
Local $bSuccess = False

$oSbOriginal = ObjCreate("Chilkat.StringBuilder")
$bSuccess = $oSbOriginal.LoadFile("qa_data/txt/utf16_emojis_accented_jap.txt","utf-16")
If ($bSuccess = False) Then
    ConsoleWrite($oSbOriginal.LastErrorText & @CRLF)
    Exit
EndIf

;  The above file contains the following text, which includes some emoji's,
;  Japanese chars, and accented chars.

$oSb = ObjCreate("Chilkat.StringBuilder")
$oSb.AppendSb($oSbOriginal)

;  Charset is not used for unicode escaping.  Set it to "utf-8", but it means nothing.
Local $sCharsetNotUsed = "utf-8"

;  Indicate the desired format/style of Unicode escaping.
;  Choose JSON-style (JavaScript-style) Unicode escape sequences by using "unicodeescape"
Local $sEncoding = "unicodeescape"

$oSb.Encode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  Output:
;  \ud83e\udde0
;  \ud83d\udd10
;  \u2705
;  \u26a0\ufe0f
;  \u274c
;  \u2713
;  \u4e2d
;  \u00e9 xyz \u00e0
;  abc \u79c1 \u306f \u3093 ghi

;  Revert back to the unescaped chars:
$oSb.Decode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  -----------------------------------------------------------------------------------------
;  Do the same, but use uppercase letters (A-F) in the hex values.
$sEncoding = "unicodeescape-upper"
$oSb.Encode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  Output:
;  \uD83E\uDDE0
;  \uD83D\uDD10
;  \u2705
;  \u26A0\uFE0F
;  \u274C
;  \u2713
;  \u4E2D
;  \u00E9 xyz \u00E0
;  abc \u79C1 \u306F \u3093 ghi

;  Revert back to the unescaped chars:
$oSb.Decode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  -----------------------------------------------------------------------------------------
;   ECMAScript (JavaScript) �code point escape� syntax

$sEncoding = "unicodeescape-curly"
$oSb.Encode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  Output:
;  \u{d83e}\u{dde0}
;  \u{d83d}\u{dd10}
;  \u{2705}
;  \u{26a0}\u{fe0f}
;  \u{274c}
;  \u{2713}
;  \u{4e2d}
;  \u{00e9} xyz \u{00e0}
;  abc \u{79c1} \u{306f} \u{3093} ghi

;  Revert back to the unescaped chars:
$oSb.Decode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  -----------------------------------------------------------------------------------------
;  Do the same, but use uppercase letters (A-F) in the hex values.
$sEncoding = "unicodeescape-curly-upper"
$oSb.Encode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  Output:
;  \u{D83E}\u{DDE0}
;  \u{D83D}\u{DD10}
;  \u{2705}
;  \u{26A0}\u{FE0F}
;  \u{274C}
;  \u{2713}
;  \u{4E2D}
;  \u{00E9} xyz \u{00E0}
;  abc \u{79C1} \u{306F} \u{3093} ghi

;  Revert back to the unescaped chars:
$oSb.Decode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  -----------------------------------------------------------------------------------------
;  HTML hexadecimal character reference

$sEncoding = "unicodeescape-htmlhex"
$oSb.Encode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  Output:
;  🧠
;  🔐
;  ✅
;  ⚠️
;  ❌
;  ✓
;  中
;  é xyz à
;  abc 私 は ん ghi

;  Revert back to the unescaped chars:
$oSb.Decode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  -----------------------------------------------------------------------------------------
;  HTML decimal character reference

$sEncoding = "unicodeescape-htmldec"
$oSb.Encode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  Output:
;  🧠
;  🔐
;  ✅
;  ⚠️
;  ❌
;  ✓
;  中
;  é xyz à
;  abc 私 は ん ghi

;  Revert back to the unescaped chars:
$oSb.Decode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  -----------------------------------------------------------------------------------------
;  Unicode code point notation or U+ notation

$sEncoding = "unicodeescape-plus"
$oSb.Encode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  Output:
;  u+1f9e0
;  u+1f510
;  u+2705
;  u+26a0u+fe0f
;  u+274c
;  u+2713
;  u+4e2d
;  u+00e9 xyz u+00e0
;  abc u+79c1 u+306f u+3093 ghi

;  Chilkat cannot unescape the Unicode code point notation or U+ notation.
;  For this style, Chilkat only goes in one direction, which is to escape.

;  To emit uppercase hex, specify unicodeescape-plus-upper
$sEncoding = "unicodeescape-plus-upper"
;  ...
;  ...

$oSb.Clear 
$oSb.AppendSb($oSbOriginal)

;  -----------------------------------------------------------------------------------------
;  Hex in Angle Brackets

$sEncoding = "unicodeescape-angle"
$oSb.Encode($sEncoding,$sCharsetNotUsed)
ConsoleWrite($oSb.GetAsString() & @CRLF)

;  Output:
;  <1f9e0>
;  <1f510>
;  <2705>
;  <26a0><fe0f>
;  <274c>
;  <2713>
;  <4e2d>
;  <e9> xyz <e0>
;  abc <79c1> <306f> <3093> ghi

;  Chilkat cannot unescape the angle bracket notation.
;  For this style, Chilkat only goes in one direction, which is to escape.

$oSb.Clear 
$oSb.AppendSb($oSbOriginal)