a$ Versus a!: Which String to Use
Rule
A PRO/5 string variable (a$) and a Java-object string variable
(a!) are not interchangeable at the byte level, even though BBj assigns
freely between them. For new code, prefer a!. In existing a$
code, stay with a$ and bring in a! only where it adds real
value, above all for Unicode work: counting, searching or slicing by character. The
input verbs INPUT, READ, READ RECORD and
ENTER fill only $-suffixed variables -- never assign one of
them directly into an a! variable.
Why
a$ functions count bytes: LEN(), CHR(),
ASC() and substring subscripts all operate byte by byte. Under UTF-8 (the
default since Java 18), a single non-ASCII character can occupy more than one byte, so
a byte count and a character count disagree the moment the string leaves plain ASCII.
a! holds a Java String, whose BBjString methods
(.length(), .indexOf(), ...) count Unicode characters, not
bytes -- the number an agent actually wants when reasoning about how many characters
a string holds.
Edge
CVS() only operates on a$ strings, and only performs
the bitwise-combinable operations in its own flag table (trim, case conversion, collapse
spaces); it has no equivalent for a!, where the same intent is a String
method such as .trim() or .toUpperCase(). Substring
subscripts (A$(pos,len)) are one-based and the second subscript is a
length, not an end position -- this only ever applies to a$.
An input verb filling a $-suffixed variable is the only legal
form: INPUT a! and READ (ch) a! are compile errors, but
READ RECORD (ch) a! passes bbjcpl and still does not work --
the compiler does not catch this one. Read into a$, then assign
a! = a$.
Evidence
Character Encoding in BBj -- the byte-versus-character distinction under UTF-8.
READ Verb,
ENTER Verb -- the input verbs that fill only
$-suffixed variables.
CVS() Function -- the operations
available only on a$.
Substrings -- the one-based, length-not-end-position subscript rule.
Types in BBj -- the
! suffix and what it marks.
Compiler result: bbjcpl -N accepts READ RECORD (ch) a!
with no error, confirming the compiler does not catch this trap.
Example
Reading into a $ variable, then converting to a!:
rem 'read into a$, then convert to a! -- never read directly into a! rec$="" READ RECORD(1)rec$ rec!=rec$
Byte count on a$ versus character count on a!:
rem 'a$ counts bytes, a! counts Java characters a$="HELLO" a!=a$ PRINT LEN(a$) PRINT a!.length()