July 2026
M	T	W	T	F	S	S
	1	2	3	4	5
6	7	8	9	10	11	12
13	14	15	16	17	18	19
20	21	22	23	24	25	26
27	28	29	30	31

Archive for the ‘UTF16’ Category

Delphi/C# reencode textfiles: some links as I thought I published sources, but didn’t

Posted by jpluimers on 2026/06/30

A long time ago (likely around 2009) I remember writing some code to re-encode text files in both Delphi and C#.

Somehow I thought I had published both, but I could only find parts of the C# code back in .NET/C# – converting UTF8 to ASCII (yes, you can loose information with this) using System.Text.Encoding.

So here are some links just in case I ever want to reproduce it in Delphi too (and a reminder to always perform this using TStream derivatives, never use TStrings or derivatives like TStringList for this):

Read the rest of this entry »

Posted in .NET, Ansi, ASCII, C#, Delphi, Development, Encoding, Mojibake, Software Development, UCS-2, UTF-16, UTF-32, UTF-8, UTF16, UTF32, UTF8, Windows-1252 | Leave a Comment »

C# Effective way to find any file’s Encoding – Stack Overflow

Posted by jpluimers on 2022/02/09

Note: notepad cannot correctly guess the encoding, see the “old new thing”: [Wayback] Some files come up strange in Notepad | The Old New Thing (talking about ANSI a.k.a. Windows-1252, UTF-16LE, UTF-16BE, UTF-8, UTF-7 somewith and some without BOM as Notepad does not understand all permutations)

David Cumps discovered that certain text files come up strange in Notepad. The reason is that Notepad has to edit files in a variety of encodings, and when its back against the wall, sometimes it’s forced to guess.

[Wayback] C# Effective way to find any file’s Encoding – Stack Overflow shows how to detect various byte order marks in C#.

–jeroen

Posted in ASCII, Development, Encoding, Software Development, Unicode, UTF-16, UTF-32, UTF-8, UTF16, UTF32, UTF8 | Leave a Comment »

PowerShell error in a script but not on the console: The string is missing the terminator: “.

Posted by jpluimers on 2021/09/29

The below one will fail in a script, both both work from the PowerShell prompt:

Success

Get-NetFirewallRule -DisplayGroup "File and Printer Sharing" | ForEach-Object { Write-Host $_.DisplayName ; Get-NetFirewallAddressFilter -AssociatedNetFirewallRule $_ }

Failure

Get-NetFirewallRule –DisplayGroup "File and Printer Sharing" | ForEach-Object { Write-Host $_.DisplayName ; Get-NetFirewallAddressFilter -AssociatedNetFirewallRule $_ }

The error you get this this:

At C:\bin\Show-File-and-Printer-Sharing-firewall-rules.ps1:5 char:52
+ ... -TCP-NoScope" | ForEach-Object { Write-Host $_.DisplayName ; Get-NetF ...
+                 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
The string is missing the terminator: ".
    + CategoryInfo          : ParserError: (:) [], ParentContainsErrorRecordException
    + FullyQualifiedErrorId : TerminatorExpectedAtEndOfString

Via [WayBack] script file ‘The string is missing the terminator: “.’ – Google Search, I quickly found these that stood out:

[WayBack] Reddit: Getting error- the string is missing the terminator: “. : PowerShell
That hyphen character in your -Reset (noticed by /u/SeeminglyScience) is an En-Dash (Decimal unicode character 8211). (Copy paste it and do [int]([char]'–') to see).

PowerShell can handle en-dash as a hyphen (for the curious, see the case for it in the PowerShell tokenizer source code here where hyphen, en-dash, em-dash and horizontalBar all fall into the same code block), but because it’s a Unicode character your source code will have to be saved as a unicode encoded .ps1 file. (UCS2-LE, UTF8+BOM, UTF8-BOM seem to work for me).

If you don’t, and you save it as ASCII/ANSI, what gets saved is â€“ which has a string opening in it. So the code becomes:
```
Set-ADAccountPassword $uname -NewPassword $newpwd â€“Reset -confirm -PassThru
```
Which now has an open string that doesn’t finish, and throws errors about unterminated strings.
Getting error- the string is missing the terminator: ".
byu/solidcore87 inPowerShell
[WayBack] azure – PowerShell script error: the string is missing the terminator: – Stack Overflow
[WayBack] powershell is missing the terminator: ” – Stack Overflow
Look closely at the two dashes in
```
unzipRelease –Src '$ReleaseFile' -Dst '$Destination'
```
This first one is not a normal dash but an en-dash (– in HTML). Replace that with the dash found before Dst.

Cause and solution

Before DisplayGroup, the first line has a minus sign and the second an en-dash. You can see this via [WayBack] What Unicode character is this ?.

Apparently, when using Unicode on the console, it does not matter if you have a minus sign (-), en-dash (–), em-dash (—) or horizontal bar (―) as dash character. You can see this in [WayBack] tokenizer.cs at function [WayBack] NextToken and [WayBack] CharTraits.cs at function [WayBack] IsChar).

When saving to a non-Unicode file, it does matter, even though it does not display as garbage in the error message.

Similarly, PowerShell has support for these special characters:

    internal static class SpecialChars
    {
        // Uncommon whitespace
        internal const char NoBreakSpace = (char)0x00a0;
        internal const char NextLine = (char)0x0085;

        // Special dashes
        internal const char EnDash = (char)0x2013;
        internal const char EmDash = (char)0x2014;
        internal const char HorizontalBar = (char)0x2015;

        // Special quotes
        internal const char QuoteSingleLeft = (char)0x2018; // left single quotation mark
        internal const char QuoteSingleRight = (char)0x2019; // right single quotation mark
        internal const char QuoteSingleBase = (char)0x201a; // single low-9 quotation mark
        internal const char QuoteReversed = (char)0x201b; // single high-reversed-9 quotation mark
        internal const char QuoteDoubleLeft = (char)0x201c; // left double quotation mark
        internal const char QuoteDoubleRight = (char)0x201d; // right double quotation mark
        internal const char QuoteLowDoubleLeft = (char)0x201E; // low double left quote used in german.
    }

The easiest solution is to use minus signs everywhere.

Another solution is to save files as Unicode UTF-8 encoding (preferred) or UTF-16 encoding (which I dislike).

–jeroen

Posted in .NET, CommandLine, Development, Encoding, PowerShell, PowerShell, Scripting, Software Development, Unicode, UTF-16, UTF-8, UTF16, UTF8 | Leave a Comment »

UTF-8 support for single byte character sets is beta in Windows and likely breaks a lot of applications not expecting this (via Unicode in Microsoft Windows: UTF-8 – Wikipedia)

Posted by jpluimers on 2018/12/04

Uh-oh: [WayBack] Unicode in Microsoft Windows: UTF-8 – Wikipedia:

Microsoft Windows has a code page designated for UTF-8, code page 65001. Prior to Windows 10 insider build 17035 (November 2017),^[7] it was impossible to set the locale code page to 65001, leaving this code page only available for:

Explicit conversion functions such as MultiByteToWideChar

The Win32 console command chcp 65001 to translate stdin/out between UTF-8 and UTF-16.

This means that “narrow” functions, in particular fopen, cannot be called with UTF-8 strings, and in fact there is no way to open all possible files using fopen no matter what the locale is set to and/or what bytes are put in the string, as none of the available locales can produce all possible UTF-16 characters.

On all modern non-Windows platforms, the string passed to fopen is effectively UTF-8. This produces an incompatibility between other platforms and Windows. The normal work-around is to add Windows-specific code to convert UTF-8 to UTF-16 using MultiByteToWideChar and call the “wide” function.^[8] Conversion is also needed even for Windows-specific api such as SetWindowText since many applications inherently have to use UTF-8 due to its use in file formats, internet protocols, and its ability to interoperate with raw arrays of bytes.

There were proposals to add new API to portable libraries such as Boost to do the necessary conversion, by adding new functions for opening and renaming files. These functions would pass filenames through unchanged on Unix, but translate them to UTF-16 on Windows.^[9] This would allow code to be “portable”, but required just as many code changes as calling the wide functions.

With insider build 17035 and the April 2018 update (nominal build 17134) for Windows 10, a “Beta: Use Unicode UTF-8 for worldwide language support” checkbox appeared for setting the locale code page to UTF-8.^[a] This allows for calling “narrow” functions, including fopen and SetWindowTextA, with UTF-8 strings. Microsoft claims this option might break some functions (a possible example is _mbsrev^[10]) as they were written to assume multibyte encodings used no more than 2 bytes per character, thus until now code pages with more bytes such as GB 18030 (cp54936) and UTF-8 could not be set as the locale.^[11]

Jump up^ [WayBack] “UTF-8 in Windows”. Stack Overflow. Retrieved July 1, 2011.

Jump up^ [WayBack] “Boost.Nowide”.

Jump up^ [WayBack] https://docs.microsoft.com/en-us/cpp/c-runtime-library/reference/strrev-wcsrev-mbsrev-mbsrev-l

Jump up^ [WayBack] “Code Page Identifiers (Windows)”. msdn.microsoft.com.

Via [WayBack] Microsoft Windows Beta UTF-8 support for Ansi API could break things. Wiki Article of the Change… – Tommi Prami – Google+

Related, as handling encoding is hard, especially if it is changed or not your default:

“windows” “17134” “utf8” – Google Search and “windows” “17134” “65001” “error” – Google Search breaks lots of applications
[RSP-21814] Microsoft Windows Beta UTF-8 support for Ansi API could break things – Embarcadero Technologies
[WayBack]

One of out customer had selected that and we started to experience very weird problems and took some time to find out why it misbehaves.

None of the application could connect to Firebird SQL server (Ours or third party) successfully.

So would be smart to go through all tooling and code with that setting, we never know what M$oft will do with that, will it ever be released or will it soon be default for all.
[WayBack] windows – Error starting SQL Server 2017 service. Error Code 3417 – Database Administrators Stack Exchange
[WayBack] windows – How do you get getLine to accept unicode characters? – Stack Overflow
[WayBack] [FIRE-23012] Windows 10’s UTF8 beta breaks URL parsing – Firestorm Bug Tracker
[WayBack] Py_Initialize: can’t initialize sys standard streams, unknown encoding: 65001 · Issue #7674 · conda/conda · GitHub
[WayBack] Blinking window while press TAB on Windows 10 · Issue #7863 · PowerShell/PowerShell · GitHub
- [WayBack] Fix dynamic changing of font in Windows console on CJK codepages by SteveL-MSFT · Pull Request #771 · lzybkr/PSReadLine · GitHub
[WayBack] cl.exe breaks with case sensitive paths – Developer Community

–jeroen

Posted in .NET, C, C++, Delphi, Development, Encoding, GB 18030, Power User, Software Development, UTF-16, UTF-32, UTF-8, UTF16, UTF32, UTF8, Windows, Windows 10 | 2 Comments »

Long read about Unicode: You, Me And The Emoji: Character Sets, Encoding And Emoji – Smashing Magazine

Posted by jpluimers on 2017/11/07

A well worth long rad:

We all recognize emoji. They’ve become the global pop stars of digital communication. But what are they, technically speaking? And what might we learn by taking a closer look at these images, characters, pictographs… whatever they are 🤔 (Thinking Face). We will dig deep to learn about how these thingamajigs work. Please note: Depending on your browser, you may not be able to see all emoji featured in this article (especially the Tifinagh characters). Also, different platforms vary in how they display emoji as well. That’s why the article always provides textual alternatives. Don’t let it discourage you from reading though! Now, let’s start with a seemingly simple question. What are emoji?

[WayBack] You, Me And The Emoji: Character Sets, Encoding And Emoji – Smashing Magazine

Via: [WayBack] Everything you ever wanted to know about characters, encodings, glyphs… and, oh yeah, emoji: bit.ly/2fNKeW3Long, rewarding read. – Ilya Grigorik – Google+

Here is just the ToC:

TABLE OF CONTENTS LINK

Character Sets And Document Encoding: An Overview

Characters

Character Sets

Coded Character Sets

Encoding

Declaring Character Sets And Document Encoding On The Web

content-type HTTP Header Declaration

Checking HTTP Headers Using A Browser’s Developer Tools

Checking HTTP Headers Using Web-based Tools

Using A Meta Element With charset Attribute

An Encoding By Any Other Name

What Were We Talking About Again? Oh Yeah, Emoji!

So What Are Emoji?

How Do We Use Emoji?

Character References

Glyphs

How Do We Know If We Have These Symbols?

The Great Emoji Proliferation Of 2016

Emoji OS Support

Emoji Support: Apple Platforms (macOS and iOS)

Emoji Support: Windows

Emoji Support: Linux

Emoji Support: Android

Emoji On The Web

Emoji One

Twemoji

Conclusion

–jeroen

Posted in ASCII, Development, Encoding, ISO-8859, ISO8859, Shift JIS, Unicode, UTF-16, UTF-8, UTF16, UTF8, Windows-1252 | Leave a Comment »

When someone writes UTF-8 and UTF-16 strings to the same file in binary format without converting between them…

Posted by jpluimers on 2017/06/21

A while ago, I had to fix some stuff in an application that would write – using a binary mechanism – UTF-8 and UTF-16 strings (part of it XML in various flavours) to the same byte stream without converting between the two encodings.

Some links that helped me investigate what was wrong, choose what encoding to use for storage and fix it:

unicode – How to Convert Ansi to UTF 8 with TXMLDocument in Delphi – Stack Overflow
delphi – What should I use? UTF8 or UTF16? – Stack Overflow
IXMLDocument.SaveToStream does not always use UTF-16 encoding | Marc Durdin’s Blog
delphi – Length() vs Sizeof() on Unicode strings – Stack Overflow
TEncoding.UTF8.GetBytes: utf 8 – String to byte array in UTF-8? – Stack Overflow
SysUtils.ByteLength() is not the nicest function, so use only for a quick-fix Delphi Unicode String Length in Bytes – Stack Overflow
AnsiString, UnicodeString et al: String Types (Delphi) – RAD Studio
Question about type identity – delphi
On the type compatibility in Delphi | The Programming Works
Joe White’s Blog » Blog Archive » Grammar details of Delphi’s “type type” feature
Type Compatibility and Identity (Delphi) – RAD Studio
Delphi “type types”: similar types but not the same type identity, some examples.
Converting in various Delphi versions (pre-Unicode and post-Unicode versions): Delphi: Encoding Strings as Python do – Stack Overflow
UTF-8 Conversion Routines – RAD Studio

–jeroen

Posted in Delphi, Delphi 10 Seattle, Delphi 10.1 Berlin (BigBen), Delphi XE8, Development, Encoding, Software Development, UTF-16, UTF-8, UTF16, UTF8, XML, XML/XSD | 3 Comments »

Some notes on stripping NULL characters and BOMs from files

Posted by jpluimers on 2017/05/31

A while ago I bumped into applications that write alternating UTF-16 and UTF-8 to files without checking what type of encoding the files were using.

So here are some notes to at least save some of the contents.

Powershell: about_Special_Characters (including NULL)
sql server – How to remove NULL char (0x00) from object within PowerShell – Stack Overflow (actually a Powershell line that does the NULL stripping)
regular expressions – Powershell 2: How to strip a specific character from a body of ASCII text – Server Fault
tr: HOWTO remove null characters from a file « Remi Bergsma’s blog

TODO: figure out how to strip the BOM.

–jeroen

Posted in Development, Encoding, Software Development, UTF-16, UTF-8, UTF16, UTF8 | Leave a Comment »

Coping with UTF-16 / UCS-2 little endian in Batch files: numbers from WMIC

Posted by jpluimers on 2016/11/22

A while ago, I needed to get the various date, time and week values from WMIC to environment variables with pre-padded zeros. I thought: easy job, just write a batch file.

Tough luck: I couldn’t get the values to expand properly. Which in the end was caused by WMIC emitting UTF-16 and the command-interpreter not expecting double-byte character sets which messed up my original batch file.

What I wanted What I got

What I wanted	What I got
`wmic_Day=21 wmic_DayOfWeek=04 wmic_Hour=15 wmic_Milliseconds=00 wmic_Minute=02 wmic_Month=05 wmic_Quarter=02 wmic_Second=22 wmic_WeekInMonth=04 wmic_Year=2015`	`Day=21 wmic_DayOfWeek=4 wmic_Hour=15 wmic_Milliseconds= wmic_Minute=4 wmic_Month=5 wmic_Quarter=2 wmic_Second=22 wmic_WeekInMonth=4 wmic_Year=2015`

wmic_Day=21
wmic_DayOfWeek=04
wmic_Hour=15
wmic_Milliseconds=00
wmic_Minute=02
wmic_Month=05
wmic_Quarter=02
wmic_Second=22
wmic_WeekInMonth=04
wmic_Year=2015

Day=21
wmic_DayOfWeek=4
wmic_Hour=15
wmic_Milliseconds=
wmic_Minute=4
wmic_Month=5
wmic_Quarter=2
wmic_Second=22
wmic_WeekInMonth=4
wmic_Year=2015

WMIC uses this encoding because the Wide versions of Windows API calls use UTF-16 (sometimes called UCS-2 as that is where UTF-16 evolved from).

As Windows uses little-endian encoding by default, the high byte (which is zero) of a UTF-16 code point with ASCII characters comes first. That messes up the command interpreter.

Lucikly rojo was of great help solving this.

His solution is centered around set /A, which:

handles integer numbers and calls them “numeric” (hinting floating point, but those are truncated to integer; one of the tricks rojo uses)
and (be careful with this as 08 and 09 are not octal numbers) uses these prefixes:
- 0 for Octal
- 0x for hexadecimal

Enjoy and shiver with the online help extract:
Read the rest of this entry »

Posted in Algorithms, Batch-Files, Development, Encoding, Floating point handling, Scripting, Software Development, UCS-2, UTF-16, UTF16 | Leave a Comment »

Some interesting encoding/Unicode/text articles on kunststube and links for test files of various encodings

Posted by jpluimers on 2016/08/17

After yesterdays post on Testing and static methods don’t go well together, I read around on Source (kunststube [WayBack]) a bit more and found these very nice articles on encoding,Unicode and text:

Related on those, some other nice readings:

Is there a set of “Lorem ipsums” files for testing character encoding issues? – Stack Overflow [WayBack]
International Components for Unicode: ICU User Guide [WayBack]
- International Components for Unicode ː Repository Browser: repos: icu/data/trunk/charset/data/ucm

ftp://ftp.unicode.org/Public/MAPPINGS [WayBack]

Notes on contents of the MAPPING directory:
EASTASIA:
    This directory is obsolete.
ETSI:
    ETSI GSM 03.38 7-bit default alphabet mapping.
ISO8859:
    These are the mapping tables of the ISO 8859 series (1 - 16).
OBSOLETE:
    Obsolete and unsupported mapping tables for historical
    and archival purposes only.
VENDORS:
    Miscellaneous mapping tables for small codesets, typically provided
    by vendors. The majority of current, useful tables are here.

–jeroen

Posted in Ansi, ASCII, CP437/OEM 437/PC-8, Development, EBCDIC, Encoding, ISO-8859, ISO8859, Shift JIS, Software Development, Unicode, UTF-16, UTF-8, UTF16, UTF8, Windows-1252 | Leave a Comment »

	A/V Revolution on Link archive: A YouTube video…
	#omdenken on Post by @lookitup.baby (Ian Co…
	xyzzy, Relay Confere… on Sad and Useless about Competit…
	ZaqHydn on MeshCore – Off grid mesh…
	ZaqHydn on MeshCore – Off grid mesh…

The Wiert Corner – irregular stream of stuff

Jeroen W. Pluimers on .NET, C#, Delphi, databases, and personal interests

Subscribe

Archives

Recent Comments

Recent Posts

Blog Stats

Meta title

Tag Cloud Title

Top Clicks

Top Posts

My badges

Twitter Updates

My Flickr Stream

Pages

All categories

Email Subscription

Archive for the ‘UTF16’ Category

Delphi/C# reencode textfiles: some links as I thought I published sources, but didn’t

C# Effective way to find any file’s Encoding – Stack Overflow

PowerShell error in a script but not on the console: The string is missing the terminator: “.

Cause and solution

UTF-8 support for single byte character sets is beta in Windows and likely breaks a lot of applications not expecting this (via Unicode in Microsoft Windows: UTF-8 – Wikipedia)

Long read about Unicode: You, Me And The Emoji: Character Sets, Encoding And Emoji – Smashing Magazine

TABLE OF CONTENTS LINK

When someone writes UTF-8 and UTF-16 strings to the same file in binary format without converting between them…

Some notes on stripping NULL characters and BOMs from files

Coping with UTF-16 / UCS-2 little endian in Batch files: numbers from WMIC

Some interesting encoding/Unicode/text articles on kunststube and links for test files of various encodings

Jeroen W. Pluimers on .NET, C#, Delphi, databases, and personal interests

Subscribe

Archives

Recent Comments

Recent Posts

Blog Stats

Meta title

Tag Cloud Title

Top Clicks

Top Posts

My badges

My Flickr Stream

Pages

All categories

Email Subscription

Archive for the ‘UTF16’ Category

Rate this:

Share this:

Rate this:

Share this:

Cause and solution

Rate this:

Share this:

Rate this:

Share this:

TABLE OF CONTENTS LINK

Rate this:

Share this:

Rate this:

Share this:

Rate this:

Share this:

Rate this:

Share this:

Rate this:

Share this: