"jf" - A portable, LLM-free interpreted programming language
* Built on a simple VM to make porting easy
* Inspired by 8-bit BASICs and Forth
* List-oriented
Parsing is inspired by Forth:
1. Read a line of input
2. Extract a word
3. Parse word as a number
4. If a number, put it on the stack
5. If not a number, look it up in the dictionary
a. If in the dictionary, parse definition
6. If not in the dictionary, throw an error
Not all words can be defined in the dictionary, there needs to be some fundamental words that are used to compose all other words. What are these "operations"? The obvious answer seems like basic arithmetic (addition, subtraction, multiplication, division) but also interactions with the hardware (moving data in memory, I/O, etc.). It seems like all hardware interaction can be reduced to memory operations if the VM maps "real" hardware to memory addresses. So something like writing text to a file becomes a memory copy operation between the memory where the text of the file currently resides into memory that is mapped to a file on disk.
Assuming that makes sense, the "foundational" words are +, -, *, / and ...hmmm
What if instead of a single stack there is just movable pointer to a memory location that is treated like a stack, but can be moved-around by the programmer as a normal part of writing a program? So for example, to draw graphics on the screen the pointer is moved to be beginning of the display memory and then numbers are entered to set the values for each pixel-representing memory location on the display. I think this can reduce the non-arithmetic words to something like BASIC's PEEK and POKE, a word to read from a specific memory location and a word to write to a specific location. If the locations are not specified, they default to an auto-incremented value based on the last operation.
To keep things simple maybe there should be only two of these pointers, one for reading and one for writing. That way the PEEK-like word has it's own pointer that will never be modified except by PEEK and the same for POKE.
-> 10 H (not really 'H', but the numeric ASCII/UTF/etc. value)
11 E
12 L
13 L
14 O
15 !
16
17
18
19
+> 20 H
21
22
23
24
25
26
27
28
29
30
-> 10 +> 20
Some implementation details: each memory location is 32 bits wide, storing up to four bytes. With this scheme multi-byte characters don't require any special handling so long as they don't exceed four bytes (I've only ever seen two). This seems wasteful, but also simple and non-character data can be packed to make sure each location is completely filled. This means memory can only be addressed at a minimum of four bytes but if everything that handles these operations is optimized for working with four byte chunks I don't think this will be a real performance problem.
So at a high level basic language features like printing "Hello, World!" to the display become:
PRINT "Hello, World!"
PRINT isn't a number, so look it up in the dictionary. PRINT is in the dictionary so begin parsing it's definition. PRINT's definition is "discard the next character, move the POKE pointer to the address of the text display buffer, then then POKE each character in turn until next character equals `"`".
There's two ways to interpret the output side of this. Each character from the input could be POKEd to the same address of the output buffer, and upon completion the buffer is flushed to the output device. Another interpretation is that the POKE pointer moves each time a new character is POKED, resulting in a range of 13 addresses populated in the output buffer requiring something to tell the system to then flush all those addresses to the display device. I lean toward the first interpretation for simplicity's sake, but this example leaves it ambiguous.
Drawing a pixel might look like this:
PIXEL 10 10 16
PIXEL's definition being "multiply the next number by the number following it, move the POKE pointer to the result then POKE the final number".
Is there much value in the default PEEK/POKE address behavior described above? It seems so intuitively, but that might be a false assumption. I'll need to noodle on that more.
What about flow control and loops?
I think both can be derived from ways to override the "instruction pointer" or whatever indicates the parser's current location in the program code. Program code is presumably stored in memory addressable by location just like the input data and memory-mapped output described above so a word like IF could be defined as "evaluate the next three words and begin parsing at the memory location named by the fourth word". In a way this is akin to how reading a word from the dictionary works; the parser's reading position is simply moved from the current input line to the dictionary line defining the word. In the case of a dictionary read the parser location returns to where it left-off in the input after the definition is parsed whereas in the case of IF parsing should continue at the destination address.
The structure and flow of IF defined this way makes conditions that match the IF's parameters an exception to the default top-down program flow. I think this is right, because the "normal" flow of a program should essentially be a list that is followed from top to bottom. Breaks from that flow should be considered exceptions, and programs that don't flow this way should be re-worked so they do.
Loops can be realized in a similar fashion by creating an IF like word that directs the parser to instructions preceding the IF. This might seem upside-down compared to other programming languages but when you think about how a person works a list, it makes sense.
1. Rinse the dishes
2. Load the dishwasher
3. Run the dishwasher
4. If the dishes are still dirty, start over
Placing the test condition for a loop before the parameters of the condition are kind of weird when you think about it.
There's details to sort-out for sure, but working these couple of examples gives me some confidence that most dictionary words could be built-up from basic math and the two PEEK and POKE-like operators.
While I'm tempted at this point to get into thinking about the VM, I'm trying to nail-down the programming language so I'll resist that temptation. The most important thing about the VM is that it is as simple as possible to make porting it to different pieces of hardware as easy as possible.
I think the best next step is to try and implement a primitive version of this and see what sorts of problems crop-up in the process.
Jason J. Gullickson, 2026