Skip to content

Lexical Structure

Grayscale source files must be encoded in UTF-8. However, identifiers are restricted to ASCII characters.

Line terminators are the ASCII line feed character (LF, U+000A) or the ASCII carriage return character (CR, U+000D) optionally followed by a line feed.

Statements and struct and enum members are separated by newlines. A semicolon ; is an optional separator equivalent to a newline, needed only to put more than one statement or member on a line:

const Point struct { x i64; y i64 }
const Color enum { RED; GREEN; BLUE }
do add(a i64, b i64) -> i64 { mut sum i64 = a + b; return sum }
do main() {
}

Two struct fields or enum variants on the same line without a ; between them are an error (E2069).

Grayscale supports two forms of comments:

Single-line comments begin with // and extend to the end of the line:

// This is a single-line comment
do main() {
mut x i64 = 42 // inline comment
}

Multi-line comments begin with /* and end with */:

do main() {
/* This is a
multi-line comment */
}

Multi-line comments can also be used inline within a statement:

do main() {
mut x i64 = /* default value */ 10
}

Multi-line comments do not nest. A /* inside a multi-line comment has no special meaning.

Identifiers name program entities such as variables, functions, types, and modules.

Identifiers must:

  • Begin with an ASCII letter or underscore
  • Contain only ASCII letters, digits, and underscores
  • Not be a reserved keyword
  • Not use the reserved prefixes gray_, _gray_, or Gray (reserved for the compiler)
  • Not be main, except as the name of the top-level entry-point function (E4026)

The standalone _ is the blank identifier (see Return Value Handling) and is not a valid variable or function name.

Valid identifiers: x, count, myVariable, point_2d, MAX_SIZE, _helper, _x

Invalid identifiers: 2fast, my-var, café

The following words are reserved and may not be used as identifiers:

Control flow:

as_long_as break case continue default
defer elif else ensure for
for_each if is loop or
or_return otherwise return switch when
while

Declarations:

alias const do enum fn import
mut new private struct use* using

use* is reserved for the import and use statement and has no other syntactic role.

Types (reserved names):

bool char Error func map
nil SourceLocation string

Error, ErrorCode, and SourceLocation are compiler-provided types (returned by error() / fallible calls and by here()), so they are always reserved. A stdlib module’s opaque type — Database, Router, Thread, Mutex, Channel, Socket, Listener, SpinLock, Arena, UUID, HttpRequest, HttpResponse — or provided enum — OpenFlag (@io), Platform (@os) — is reserved only while that module is imported; otherwise the name is free for a user struct or enum.

Sized types (reserved names):

i8 i16 i32 i64 i128 i256
u8 u16 u32 u64 u128 u256
f32 f64

Operators and values:

bit_and bit_not bit_or bit_shift_left bit_shift_right
bit_xor cast false in not_in range
true

Some keywords have shorter or more familiar aliases. Both forms are identical and produce the same token:

Alias Canonical Purpose
else otherwise Default branch
elif or Else-if branch
while as_long_as Condition loop
fn do Function declaration
switch when Pattern match
case is Pattern match branch
defer ensure Deferred call
!in not_in Non-membership test

Each pair is tracked independently, so one file may write while and fn while another writes as_long_as and do. Mixing the two spellings of a single pair within one file is an E2088 error.

Joint pairs. Two of the pairs span two keywords each, and both words move together — a file must take both from the same side or neither:

Dialect Pattern match Branch chain
Canonical when … is or … otherwise
Alias switch … case elif … else
do main() {
mut x i64 = 1
mut a bool = true
mut b bool = false
switch x { case 1 { } default { } } // ok
// when x { is 1 { } default { } } // ok, but not in the same file as switch
// switch x { is 1 { } default { } } // E2088 — crossed dialects
if a { } elif b { } else { } // ok
// if a { } or b { } otherwise { } // ok, but not in the same file as elif/else
// if a { } elif b { } otherwise { } // E2088 — crossed dialects
}

if and default are spelled the same in both dialects and never vary.

+ - * / %
== != < > <= >=
&& || !
= += -= *= /= %=
++ --
^ & @ #
( ) { } [ ]
, : . ->
  • ^ — pointer type prefix and dereference postfix (^Type, ptr^)
  • & — mutable parameter marker in function signatures
  • @ — module prefix in imports (import @math)
  • # — attribute prefix (#doc, #flags, #strict)

Bitwise operations use keyword syntax (bit_and, bit_or, etc.) because ^ and & are already used for pointer types and mutable parameters. See Bitwise Operators.

Integer literals represent integer values.

Underscores may be used for readability but:

  • Must not appear at the beginning or end
  • Must not appear consecutively
  • Must not appear adjacent to the decimal point

Examples: 42, 1_000_000, 0xFF, 0xDEAD_BEEF, 0o777, 0o1_2_3, 0b1010, 0b1111_0000

Floating-point literals represent floating-point values.

Examples: 3.14159, 0.5, 100.0

String literals represent string values.

Escape sequences:

  • \n - line feed (U+000A)
  • \r - carriage return (U+000D)
  • \t - horizontal tab (U+0009)
  • \a - bell (U+0007)
  • \b - backspace (U+0008)
  • \f - form feed (U+000C)
  • \v - vertical tab (U+000B)
  • \0 - NUL byte (U+0000)
  • \\ - backslash
  • \" - double quote
  • \' - single quote
  • \$ - literal $ (suppresses string interpolation, e.g. "\${x}" is the text ${x})
  • \xNN - hex byte value (e.g., \x48 = ‘H’, \x0a = newline)

String interpolation allows embedding expressions within strings:

do main() {
mut name string = "World"
mut greeting string = "Hello, ${name}!" // "Hello, World!"
}

Strings may also be joined with the + operator ("Hello, " + name + "!"); see Arithmetic Operators.

Raw string literals are enclosed in backticks and do not process escape sequences or string interpolation:

Raw strings:

  • Do not process escape sequences (\n is a literal backslash followed by n)
  • Do not process string interpolation (${x} is literal text)
  • May span multiple lines
  • Cannot contain backticks (no escape mechanism)
do main() {
mut path string = `C:\Users\test\file.txt`
mut pattern string = `\d+\.\d+`
mut multi string = `line1
line2
line3`
}

A character literal is one Unicode codepoint between single quotes — one character or one escape sequence. Its type is char.

do main() {
'A' '7' ' ' 'é' '日' '\n' '\u{1F600}'
}

Escape sequences:

Escape Meaning
\n \r \t LF, CR, tab
\\ \' \" backslash, single quote, double quote
\0 NUL (U+0000)
\xNN codepoint U+00NN — exactly two hex digits
\u{H…} codepoint from 1–6 hex digits (U+0000–U+10FFFF)
  • A character literal must contain exactly one codepoint. '' and 'ab' are E1018.
  • The bytes between the quotes are decoded as UTF-8; a malformed sequence is E1018.
  • In a character literal \xNN is codepoint U+00NN, not a raw byte (unlike a string literal, where \xNN is a byte). For sub-codepoint byte values use u8.
  • An unterminated literal is E1005; an unknown escape is E1007; a malformed \x or \u{} is E1006.

true and false are the two boolean literals.

The literal nil represents the absence of a value.