a$ Versus a!: Which String to Use

Rule

A PRO/5 string variable (a$) and a Java-object string variable (a!) are not interchangeable at the byte level, even though BBj assigns freely between them. For new code, prefer a!. In existing a$ code, stay with a$ and bring in a! only where it adds real value, above all for Unicode work: counting, searching or slicing by character. The input verbs INPUT, READ, READ RECORD and ENTER fill only $-suffixed variables -- never assign one of them directly into an a! variable.

Why

a$ functions count bytes: LEN(), CHR(), ASC() and substring subscripts all operate byte by byte. Under UTF-8 (the default since Java 18), a single non-ASCII character can occupy more than one byte, so a byte count and a character count disagree the moment the string leaves plain ASCII. a! holds a Java String, whose BBjString methods (.length(), .indexOf(), ...) count Unicode characters, not bytes -- the number an agent actually wants when reasoning about how many characters a string holds.

Edge

CVS() only operates on a$ strings, and only performs the bitwise-combinable operations in its own flag table (trim, case conversion, collapse spaces); it has no equivalent for a!, where the same intent is a String method such as .trim() or .toUpperCase(). Substring subscripts (A$(pos,len)) are one-based and the second subscript is a length, not an end position -- this only ever applies to a$.

An input verb filling a $-suffixed variable is the only legal form: INPUT a! and READ (ch) a! are compile errors, but READ RECORD (ch) a! passes bbjcpl and still does not work -- the compiler does not catch this one. Read into a$, then assign a! = a$.

Evidence

Character Encoding in BBj -- the byte-versus-character distinction under UTF-8.

READ Verb, ENTER Verb -- the input verbs that fill only $-suffixed variables.

CVS() Function -- the operations available only on a$.

Substrings -- the one-based, length-not-end-position subscript rule.

Types in BBj -- the ! suffix and what it marks.

Compiler result: bbjcpl -N accepts READ RECORD (ch) a! with no error, confirming the compiler does not catch this trap.

Example

Reading into a $ variable, then converting to a!:

rem 'read into a$, then convert to a! -- never read directly into a!
rec$=""
READ RECORD(1)rec$
rec!=rec$

Byte count on a$ versus character count on a!:

rem 'a$ counts bytes, a! counts Java characters
a$="HELLO"
a!=a$
PRINT LEN(a$)
PRINT a!.length()