# mcpu-cpp `mcpu-cpp` is the preprocessor for MCPU programming languages. It is an independent component of the LibMPU/LibMPUIO/LibMCPU ecosystem and is not tied to the name of any single language: the active language is selected with the `#lang` directive. This document defines the normative behavior of `mcpu-cpp`: its text model, directives, macro engine, include pipeline, configuration, diagnostics, and dependency generation for MCPU tools. ## 1. Text model External source files and configuration files are encoded in UTF-8. The UTF-8 must be valid. For source programs, the check that characters belong to the UCS-2 range is performed after comments have been removed. Therefore a valid Unicode scalar value above `U+FFFF` is permitted inside a comment, but remains an error in program text. After this stage, source text is processed as a sequence of `__mpu_char16_t` values. An input UTF-8 BOM is accepted and removed. An embedded NUL in a source file is forbidden. `CRLF` and `CR` line endings are normalized to `LF`. ## 2. Actions performed independently of directives `mcpu-cpp` performs several transformations before directives are parsed. ### 2.1. Backslash-newline A `\\` immediately followed by a newline is removed before comments, directives, and macros are recognized. For example, ```text #defi\ ne FOO 10\ 20 ``` is equivalent to the logical line ```text #define FOO 1020 ``` Physical line numbers continue to contribute to the current source position. Unless the user changes that position with `#line`, those physical positions are the ones reflected in generated line markers. ### 2.2. Comments `/* ... */` and `// ...` comments are removed before subsequent processing. Where needed to keep adjacent tokens separate, a whitespace separator is preserved. If a comment terminates a nonempty line, neither a synthetic separator nor whitespace that immediately preceded the comment is retained after the comment is removed: the line ends at its last significant character. The same rule applies to a multi-line comment that starts after program text. If comment removal leaves a line containing only whitespace, the line becomes genuinely empty. A comment between two tokens still leaves the separator required to keep the tokens from being joined. Newlines are preserved so source coordinates are not destroyed. Comments are not recognized inside string or character constants. In the `diff` language, an apostrophe is not treated as the beginning of a character constant because it is used in derivative notation. Within a literal `#include <...>` operand, `/*` and `//` sequences are treated as part of the file name. ## 3. Directives and the output stream A directive begins with `#` when only whitespace or comments precede it on the logical line. Whitespace is permitted between `#` and the directive name. Source-position service information in the output stream uses GNU **line markers**: ```text # line-number "file-name" [flags] ``` This is not the input directive `#line`. Entering an included file adds flag `1` to the line marker, and returning to the file that contained the `#include` adds flag `2`. These values have the same meaning as in GNU CPP: `1` means entering a new file and `2` means returning to the previous file. Flag `2` is not a nesting count or include level. For example: ```text # 1 "main.c" # 1 "defs.h" 1 ... # 2 "main.c" 2 ``` The input directive ```text #line 62 "main.y" ``` is not copied to the output stream. It changes the logical values of `__LINE__` and `__FILE__` for subsequent text and is represented in output by a line marker: ```text # 62 "main.y" ``` The arguments of `#line` undergo macro expansion according to the line-control model. If an `#include` follows such a `#line`, the return marker receives flag `2`, for example `# 65 "main.y" 2`. A name installed by `#line` becomes the logical name used by `__FILE__` and line markers; it does not change the directory used to resolve a quoted `#include`. Preprocessor directives use canonical English names only. Unicode remains fully supported in identifiers, strings, comments, and other user text. ## 4. Header files The following forms are supported: ```text #include "file" #include #include_next "file" #include_next #pragma once ``` For ordinary `#include "file"`, the directory of the **physical** current source file is always checked first. A logical name established by `#line` does not affect this step. For `#include `, the directory of the current file is not checked. ### 4.1. Relocatable MCPU root as an ecosystem-wide principle Starting with release 0.0.37, the MCPU installation directory **does not contain a version number of a particular tool** and is not an absolute runtime constant compiled into the binary. A version belongs to `mcpu-cpp`, `mcpu-as`, `mcpu-ld`, `mcpu-run`, or a library; it does not define the root of the shared MCPU environment. For a typical configuration: ```text ./configure --prefix=/usr --libdir=/usr/lib64 ``` `make install` creates: ```text /usr/lib64/mcpu/ ├── bin/ │ └── mcpu-cpp ├── etc/ │ └── mcpu-cpp.conf ├── include/ │ ├── diff/ │ ├── dift/ │ ├── alg/ │ ├── as/ │ ├── avm/ │ └── acs/ └── lib/ # common directory for future MCPU libraries ``` The public program name lives in `$bindir`: ```text /usr/bin/mcpu-cpp -> ../lib64/mcpu/bin/mcpu-cpp ``` The absolute `/usr/lib64/mcpu` path is **not part of the MCPU-CPP runtime ABI**. It is only the configure-time installation location selected by `make install`. On every normal invocation, MCPU-CPP determines the actual path of its own executable through Linux `/proc/self/exe`. The public-command symlink does not interfere with this: `/proc/self/exe` names the binary that is actually being executed. If `/proc/self/exe` is unavailable, a fallback resolves `argv[0]` through `PATH` and `realpath(3)`; there is no fallback to a compiled-in configure-time installation root. For an executable ```text /bin/mcpu-cpp ``` the runtime root is derived as: ```text executable = /bin/mcpu-cpp executable dir = /bin MCPU runtime root = ``` and the following paths are derived from it automatically: ```text /etc/mcpu-cpp.conf /include ``` Therefore the whole tree can be physically moved, for example from ```text /usr/lib64/mcpu/ ``` to ```text /opt/mcpu-test/ ``` or ```text $HOME/devel/mcpu-next/ ``` and `/bin/mcpu-cpp` immediately starts using `/etc/mcpu-cpp.conf` and `/include` without being reconfigured. The old absolute path is retained neither in runtime defaults nor in the installed `mcpu-cpp.conf`. This is not a preprocessor-specific trick; it is a **general MCPU ecosystem principle**. Future `mcpu-as`, `mcpu-ld`, `mcpu-run`, libraries, CRT, and other components are expected to share one relocatable root: ```text /bin /etc /include /lib ``` Their own versions may differ. Consistency of a particular MCPU environment is defined by all components residing in one runtime tree, not by matching version suffixes in directory names. ### 4.2. Runtime defaults, configuration layers, and the system include root Before reading any configuration file, MCPU-CPP creates the runtime-derived value: ```text MCPU_CPP_SYSTEM_INCLUDE_PATH = /include ``` Configuration layers are then applied in order of increasing priority: ```text runtime-derived defaults ↓ /etc/mcpu-cpp.conf ↓ /etc/mcpu/mcpu-cpp.conf ↓ $HOME/.mcpu/etc/mcpu-cpp.conf ``` `/etc/mcpu-cpp.conf` is installed with MCPU-CPP, but deliberately does not contain an absolute default `MCPU_CPP_SYSTEM_INCLUDE_PATH`: otherwise moving the tree would restore the old path. `/etc/mcpu/mcpu-cpp.conf` is an optional machine-wide override; `make install` does not create `/etc/mcpu`. `$HOME/.mcpu/etc/mcpu-cpp.conf` is also optional, is not versioned, and has the highest configuration priority. If one variable is defined more than once, the last definition wins, including an empty definition. Therefore `MCPU_CPP_SYSTEM_INCLUDE_PATH` remains a fully replaceable system root. For example: ```text MCPU_CPP_SYSTEM_INCLUDE_PATH = $HOME/mcpu-next/include; ``` completely replaces the runtime-derived `/include`. For active `#lang "as"`, the following locations are then checked: ```text $HOME/mcpu-next/include/as $HOME/mcpu-next/include ``` Standard language subdirectories are always derived by the preprocessor from one root; there are no variables named `MCPU_CPP_SYSTEM__INCLUDE_PATH`. An empty effective value: ```text MCPU_CPP_SYSTEM_INCLUDE_PATH = ; ``` removes the configured system stage entirely. A higher-priority configuration file may later enable it again with a nonempty value. `--config-file FILE` applies an explicitly selected file on top of the runtime-derived default. `--no-config` disables **only configuration-file reading**: `/etc/mcpu-cpp.conf`, `/etc/mcpu/mcpu-cpp.conf`, and `$HOME/.mcpu/etc/mcpu-cpp.conf` are not read, but `/include` remains the standard system root. `--sys-root=PATH` is a stronger command-line override. It uses `PATH` as the MCPU system root, makes `PATH/include` the runtime-derived `MCPU_CPP_SYSTEM_INCLUDE_PATH`, and implicitly disables all configuration-file reading, including an explicitly supplied `--config-file`. `PATH` may be absolute or relative. A relative value is resolved against the invocation working directory before the effective include root is formed. This makes it possible to select another complete MCPU tree for one invocation without changing either the installed tree or persistent configuration. Only `-nostdinc` removes the effective standard system tree from include search for one invocation; an explicit `-isystem` still remains a command-line directory. ### 4.3. Normative include-file search order Search order is part of the MCPU-CPP contract. Explicit command-line parameters have priority over persistent configuration. After the optional directory of the current physical file, the effective chain is strictly: ```text explicit -I ↓ explicit -isystem ↓ MCPU_CPP__INCLUDE_PATH ↓ MCPU_CPP_INCLUDE_PATH ↓ MCPU_CPP_SYSTEM_INCLUDE_PATH/ ↓ MCPU_CPP_SYSTEM_INCLUDE_PATH ↓ explicit -idirafter ↓ MCPU_CPP_AFTER_INCLUDE_PATH ``` Entries that are absent or do not contain the requested file are skipped. `MCPU_CPP__INCLUDE_PATH` denotes user-configurable language-specific path lists: ```text MCPU_CPP_DIFF_INCLUDE_PATH MCPU_CPP_DIFT_INCLUDE_PATH MCPU_CPP_ALG_INCLUDE_PATH MCPU_CPP_AS_INCLUDE_PATH MCPU_CPP_AVM_INCLUDE_PATH MCPU_CPP_ACS_INCLUDE_PATH ``` The user fully controls the names and locations of these directories. `MCPU_CPP_INCLUDE_PATH` is a common user path list visible in every language state. `-idirafter` and `MCPU_CPP_AFTER_INCLUDE_PATH` form a common fallback area. MCPU-CPP does not automatically derive `` subdirectories for them. The user controls their internal layout and may, for example, write: ```text #include ``` Priority is determined by the semantic class, not by the relative appearance of different classes in argv or configuration. Within one class, insertion order is preserved. ### 4.4. `#include_next` and wrapper headers `#include_next` is intended primarily for wrapper headers. It allows a local header to precede a system header, adjust local policy, and then continue the search for a same-named header along the normative chain without copying the system file or using an absolute name. For example: ```text mcpu-cpp -isystem $HOME/mcpu-wrapper ... ``` with `$HOME/mcpu-wrapper/math.h`: ```text #ifndef SOME_SYSTEM_MACRO #define SOME_SYSTEM_MACRO temporary_value #define REMOVE_SOME_SYSTEM_MACRO 1 #endif #include_next #ifdef REMOVE_SOME_SYSTEM_MACRO #undef SOME_SYSTEM_MACRO #undef REMOVE_SOME_SYSTEM_MACRO #endif ``` If the home configuration also specifies: ```text MCPU_CPP_SYSTEM_INCLUDE_PATH = $HOME/mcpu-next/include; ``` a wrapper found through `-isystem` continues `#include_next` through the configured user paths, then through `$HOME/mcpu-next/include/` and `$HOME/mcpu-next/include`. The old system tree at the original installation location does not participate. This is the intended way for a system developer or tester to work in a private sandbox. MCPU-CPP stores the exact **physical element of the effective search chain** from which the current header was found. `#include_next` starts at the next element. The `"file"` and `` forms of `#include_next` are equivalent; the directory of the current file is not checked again. If the current file was found by ordinary quoted search relative to its containing file and therefore has no search-chain provenance, `#include_next` starts at the first element of the configured chain. The operand may be produced by macro expansion. A logical name installed by `#line` does not affect physical provenance. If no suitable file exists after the current entry, preprocessing fails. ### 4.5. `#pragma once` An active ```text #pragma once ``` directive marks the **physical file** as already processed during the current MCPU-CPP invocation. A later attempt to include the same physical file skips its contents. The directive itself is consumed by the preprocessor and is not copied to output, including in `-dD` mode. Identity is determined by the file-system `st_dev`/`st_ino` pair, not by the path string. Therefore the same file cannot bypass `#pragma once` by being reached as `./file.h`, through a symbolic link, or through another hard-link name. A logical name installed by `#line` also has no effect on this physical identity. The mark takes effect immediately when the active directive is processed. Therefore a header may include itself after `#pragma once`: the repeated include is skipped and recursion does not occur. A directive in an inactive conditional branch has no effect. MCPU-CPP recognizes only the exact `#pragma once` form, with optional whitespace. Other `#pragma` directives are not interpreted by the preprocessor and are preserved for later compiler stages; for example, `#pragma pack(...)` continues to be passed through to output. `#pragma once` supplements, but does not modify, the normative `#include`/`#include_next` search chain. The ordinary search mechanism first finds a physical file, then the `once` registry decides whether its contents must be processed. ### 4.6. Forced files: `-imacros FILE` and `-include FILE` The command-line options ```text -imacros FILE -include FILE ``` process a file before the primary input. They use the ordinary preprocessing engine, not a separate simplified parser. The normative start-of-translation-unit order is: ```text predefined macros -> -D/-U in command-line order -> all -imacros in command-line order -> all -include in command-line order -> primary input ``` Thus the relative interleaving of `-imacros` and `-include` in `argv` does not interleave the two groups: **all** `-imacros` files are always processed before **all** `-include` files. `-imacros FILE` fully preprocesses the file. Its `#define`/`#undef`, conditional directives, `#lang`/`#endlang`, `#include`, `#include_next`, `#pragma once`, and diagnostics have normal semantics. However, all normal preprocessing output from this forced file, including line markers and text from nested headers, is discarded. The resulting macro-table state and other preprocessing state are retained for later forced files and for the primary input. `-include FILE` uses the same machinery, but preserves normal output, as if the located header had been included immediately before the primary source. A forced include is a real include boundary: inside it `__INCLUDE_LEVEL__ == 1`, inside a header that it includes the level is `2`, and the primary input remains at level `0`. `__BASE_FILE__` inside forced files remains the name of the primary input. An absolute forced-file operand is used directly. A relative operand is first searched for in the **current working directory**, then along the ordinary include chain: ```text explicit -I explicit -isystem MCPU_CPP__INCLUDE_PATH MCPU_CPP_INCLUDE_PATH MCPU_CPP_SYSTEM_INCLUDE_PATH/ MCPU_CPP_SYSTEM_INCLUDE_PATH explicit -idirafter MCPU_CPP_AFTER_INCLUDE_PATH ``` The directory of the primary input receives no special priority when resolving an `-imacros`/`-include` operand. Once a forced file is found, ordinary quoted `#include "file"` inside it is again resolved relative to the physical directory of that forced file. If the forced file was found through an element of the include chain, its provenance is retained and `#include_next` continues at the next chain element. Forced files and the headers actually reached from them participate in the ordinary physical dependency registry. Their user/system classification is derived from the same search provenance, so `-MM`/`-MMD` filter system forced headers exactly as they filter ordinary system headers. A missing forced file is an error. ### 4.7. Dependency generation: `-M`, `-MM`, `-MG`, `-MD`, `-MMD`, `-MF`, `-MT`, `-MQ` `-M` and `-MM` use **the same include-pipeline pass** as ordinary preprocessing. There is no second, independent header search. The dependency graph therefore inherits the normative search order, `#include_next`, macro-expanded include operands, conditional compilation, and `#pragma once` automatically. `-M` suppresses normal preprocessing output and emits one Make rule: ```make file.o: file.c header1.h header2.h ``` The list contains the primary source and all physical headers actually reached, including system headers. A physical file appears only once. Identity is `st_dev + st_ino`, so an alternate relative spelling, symbolic link, or hard link does not create a duplicate dependency. A logical name established by `#line` is only a source name and never enters the dependency list. `-MM` builds the same graph but removes system dependencies. System context includes headers found through explicit `-isystem`, the configured system tree `MCPU_CPP_SYSTEM_INCLUDE_PATH/` / `MCPU_CPP_SYSTEM_INCLUDE_PATH`, explicit `-idirafter`, and `MCPU_CPP_AFTER_INCLUDE_PATH`, as well as the entire branch of headers included directly or indirectly from such a system header. The syntax `#include "file"` versus `#include ` does not by itself determine whether a dependency is a system dependency. If one physical file is reached from a system branch and is later included directly from user context, it remains a user dependency and is present in `-MM` output. The default target is derived from the basename of the primary source: its suffix is replaced with the object suffix (`.o` by default). Paths and the default target are Make-quoted. For stdin, the GNU-like form is `-: -`. `-MD` and `-MMD` use the same dependency graph but, unlike `-M` and `-MM`, **do not suppress normal preprocessing output**. `-MD` includes system headers like `-M`; `-MMD` applies the user-only filter of `-MM`. One pass can therefore produce both preprocessed text and a side-effect dependency file. If `-MF` is not specified, side-effect mode chooses the `.d` file name automatically: * without `-o`, the input basename loses its suffix and receives `.d`; input pathname directories are not copied into the dependency-file name; * with an ordinary `-o FILE`, the output-file suffix is replaced by `.d`; * stdin uses `-.d`. `-MF FILE` overrides the automatic dependency-file name. `-MF -` means stdout. `-MF` also works with dependency-only `-M`/`-MM`; in that case it takes priority over the ordinary destination of the Make rule. `-MF` by itself, without one of `-M`, `-MM`, `-MD`, or `-MMD`, is an error. The semantics intentionally follow GNU CPP: `-MD`/`-MMD` do not accept their own argument; `-MF` is a separate dependency-output option. `-MT TARGET` replaces the automatic target with `TARGET` **exactly as supplied**. No Make quoting is performed. Thus one `-MT` argument may contain multiple targets separated by spaces: ```text -MT 'obj/a.o obj/a.pic.o' ``` and repeated `-MT` options also append targets to the same rule: ```text -MT obj/a.o -MT obj/a.pic.o ``` `-MQ TARGET` has the same target-selection semantics but quotes characters that are special to Make. For example: ```text -MQ '$(OBJDIR)/foo.o' ``` produces the left-hand side: ```make $$(OBJDIR)/foo.o: ``` Both separate arguments (`-MT TARGET`, `-MQ TARGET`) and attached forms (`-MTTARGET`, `-MQTARGET`) are supported. If at least one `-MT` or `-MQ` is present, the automatic default target is not emitted. In particular, `--object-suffix` affects only the automatic target and does not rewrite explicit targets. When no explicit target is present, the default target is Make-quoted as with `-MQ`. Repeated and mixed `-MT`/`-MQ` options are allowed. As in GNU CPP, all `-MT` targets are emitted first in their command-line order, followed by all `-MQ` targets in their command-line order. All of them form the left-hand side of **one** dependency rule. `-MT` and `-MQ` are meaningful only with one of `-M`, `-MM`, `-MD`, or `-MMD`. Using either without dependency generation is a command-line error. `-MG` changes only the handling of **missing** include files in dependency-only `-M` and `-MM` modes. Without `-MG`, an unresolved `#include` remains an error. With `-M -MG` or `-MM -MG`, a missing header is treated as a future generated file: preprocessing does not fail and the directive operand is added to the dependency rule **exactly as obtained after macro expansion**, without prepending a guessed include directory. For example: ```text #include "generated.h" ``` adds `generated.h` under `-M -MG`, even when that file does not yet exist. A macro-expanded include behaves the same way: the dependency receives the expanded name. `-MG` is valid only with `-M` or `-MM`; combinations with `-MD`/`-MMD`, or use without dependency-only mode, are command-line errors. Unresolved dependencies are integrated into **the same ordered dependency registry** as physical files, but occupy a separate identity domain. For a found file, the registry still uses `st_dev/st_ino` and physical provenance. For a missing file there is no such information, so an `-MG` entry performs no `stat()` and is deduplicated by the exact include-operand text. This is essential: an identically named file in the current working directory must not turn an unresolved `` into a false physical match when angle search did not find that file. Different unresolved spellings, such as `generated.h` and `./generated.h`, are distinct dependencies. For `-MM`, an unresolved dependency gets its user/system class from search context: a missing `` is system-class and a missing `"file"` is user-class when the including unit is not itself a system header; any missing include reached from a system header remains system-class. If the same unresolved operand appears more than once, the classification of its first occurrence is retained, matching GNU CPP. Physical dependencies keep the existing rule that if the same inode is later reached from user context, it is no longer system-only. `-MG` also applies to missing command-line forced files `-include FILE` and `-imacros FILE`: their operand enters the unresolved registry as a user dependency without a synthetic search prefix. When the file exists, `-include`/`-imacros` continue to use the normal physical dependency registry and search provenance. ## 5. Language switching The preprocessor starts in language state `0`. This is the unnamed primary C-like language and it is not a valid argument of `#lang`. Supported languages are: | Name | Purpose | |---|---| | `diff` | differential equations | | `dift` | difference equations | | `alg` | algebraic equations | | `as` | MCPU assembler (`mcpu-as`) | | `avm` | analog-computer schemes | | `ACS` | block diagrams of automatic-control systems | `#lang` must be followed by a string constant containing one nonempty word: ```text #lang "diff" ``` The name is checked only against the internal language list above and is compared case-insensitively in ASCII. Thus `"diff"`, `"Diff"`, `"DIFF"`, and `"dIfF"` select the same language. The spelling inside the quotes is preserved in output. Whitespace inside the string constant is forbidden: `" diff"`, `"diff "`, and `"di ff"` are errors. Escape sequences are not interpreted inside this constant. The closing quote must occur on the same physical source line. Only whitespace is permitted between the closing quote and the end of the line. Whitespace outside the string is normalized. For example: ```text # lang "DiFf" ``` becomes: ```text #lang "DiFf" ``` `#lang` pushes a new language onto the language stack; `#endlang` restores the previous language. The stack is not reset by `#include`, so a language block may begin and end in different files. `#lang` and `#endlang` remain in the output stream for the later frontend dispatcher; `#lang` is emitted in normalized form. ## 6. Object-like macro definitions Starting with 0.0.4, object-like macros are supported: ```text #define BUFFER_SIZE 1024 #define NAME value #define EMPTY ``` The `#define` directive itself is not copied to normal output. In ordinary text, a macro identifier is replaced with its replacement list. That replacement is rescanned for macro names, so cascaded expansion works: ```text #define A B #define B 10 A ``` produces `10`. While a particular macro is being expanded, that macro is temporarily disabled. Therefore self-referential and mutually recursive definitions do not cause infinite recursion. Macro names are not expanded inside string or character constants. For the `diff` language, an apostrophe keeps its language-specific meaning and does not protect following text as a C character constant. A multi-line definition using backslash-newline is supported because splicing occurs before `#define` is parsed. ### 6.1. `#undef` ```text #undef NAME ``` removes an object-like macro definition. Undefining a nonexistent macro is not an error. ### 6.2. Computed `#include` An `#include` argument that does not begin directly with `"` or `<` is first macro-expanded. Therefore both of these are valid: ```text #define HEADER #include HEADER ``` and ```text #define HEADER "local.h" #include HEADER ``` The expansion result must have the form `"file"` or ``. ## 7. Function-like macros The macro engine supports function-like macros: ```text #define identifier( argument-list ) replacement ``` The opening parenthesis in the definition must follow the macro name **immediately**. Thus ```text #define F(X) X ``` defines a function-like macro, while ```text #define F (X) ``` defines an object-like macro with replacement `(X)`. At a use site, whitespace is permitted between a function-like macro name and the opening parenthesis. If `(` does not follow, the identifier is not a call of that macro and remains in output. For an ordinary function-like macro, the number of actual arguments must equal the number of formal parameters. For a variadic macro, every fixed argument must be present, while the variadic tail may contain any number of arguments, including an empty tail. Nested parentheses are tracked while parsing actual arguments; a comma inside them does not separate arguments. Square brackets do not have this property; this is part of the adopted macro-expansion semantics. For example: ```text #define min(X, Y) ((X) < (Y) ? (X) : (Y)) min(1, 2) ``` produces: ```text ((1) < (2) ? (1) : (2)) ``` Before substitution, an ordinary actual argument itself undergoes macro expansion. Cascaded and nested calls therefore work naturally: ```text #define A 7 #define min(X, Y) ((X) < (Y) ? (X) : (Y)) min(min(A, 3), 10) ``` A formal parameter may occur any number of times in the replacement list. An expression with side effects in an actual argument may therefore be evaluated multiple times by the later compiler; the preprocessor does not attempt to repair such source code. Macros with no formal parameters are supported: ```text #define READY() 1 ``` They expand only when called as `READY()` (whitespace between the name and `(` is permitted at a use site); the standalone identifier `READY` does not expand. Formal parameter names must be distinct. An unterminated parameter list, invalid punctuation, and too few or too many actual arguments are errors. ### 7.1. Stringification `#` The stringification operator (`#`) is supported for parameters of function-like macros: ```text #define STR(X) #X STR(alpha + beta) ``` produces: ```text "alpha + beta" ``` Stringification uses the **raw actual argument before macro expansion**. Thus: ```text #define A 7 #define STR(X) #X #define XSTR(X) STR(X) STR(A) -> "A" XSTR(A) -> "7" ``` Leading and trailing whitespace in the argument is removed. Internal whitespace sequences are collapsed to one space except inside string/character tokens of the active language. Double quotes and backslashes inside quoted tokens are escaped so that the result remains one valid string constant. In a function-like replacement list, `#` must refer, directly or after whitespace, to a formal parameter name. Inside a quoted token, `#` is not an operator. An empty actual argument is allowed and stringifies as `""`. ### 7.2. Token concatenation `##` Starting with 0.0.21, token concatenation (`##`) is supported with semantics aligned with GNU CPP and the macro engine's `collect_expansion()` / `macroexpand()` model. The operator combines two adjacent preprocessing tokens into one token, after which the resulting replacement list is rescanned for macro expansion. For example: ```text #define CAT(A, B) A ## B CAT(foo, bar) ``` produces `foobar`. Concatenation may form an identifier, preprocessing number, or multi-character punctuator. For example: ```text CAT(1.5, e3) -> 1.5e3 CAT(+, =) -> += ``` When a formal parameter is directly adjacent to `##`, its actual argument is substituted **without preliminary macro expansion**. This is the same raw argument principle used by stringification. To expand first and concatenate second, use the normal two-level GNU CPP pattern: ```text #define AFTERX(X) X_ ## X #define XAFTERX(X) AFTERX(X) #define TABLESIZE 1024 #define BUFSIZE TABLESIZE AFTERX(BUFSIZE) -> X_BUFSIZE XAFTERX(BUFSIZE) -> X_1024 ``` An empty actual argument adjacent to `##` behaves as a placemarker: it adds no token, and concatenation on that side leaves the remaining operand unchanged. If an actual argument contains multiple preprocessing tokens, only the edge token directly adjacent to `##` is concatenated; the others are preserved and participate in the subsequent rescan. `#` and `##` may be used in the same function-like macro, for example: ```text #define COMMAND(NAME) #NAME | NAME ## _command ``` Here `#NAME` uses the raw spelling of the argument for stringification, while `NAME ## _command` uses the same raw argument for concatenation. Inside a quoted token, `##` is not an operator. Comments have already become whitespace by the time macro expansion occurs, so comments cannot be created by concatenating `/` and `*`. Whitespace may originally appear between `##` and its operands; it does not participate in the concatenation. If the two operands do not form one valid preprocessing token, a diagnostic is issued and the original tokens are retained; whether whitespace appears between them after that diagnostic is not part of the contract. `##` at the beginning or end of a replacement list is a macro-definition error. ### 7.3. Variadic macros: `...` and `__VA_ARGS__` Starting with 0.0.46, variadic function-like macros are supported in the modern C99-compatible form: ```text #define LOG(...) output(__VA_ARGS__) #define LOGF(format, ...) output(format, __VA_ARGS__) ``` The `...` marker may be the only parameter or the final element after one or more fixed parameters. The old GNU extension with a named variadic parameter, ```text #define LOG(args...) ... ``` is intentionally not supported in 0.0.46. `__VA_OPT__` was also not part of that particular release. At invocation, every token after the last fixed parameter, including commas that separate those tokens, forms one logical variable argument and is substituted for `__VA_ARGS__`. In an ordinary position, that variable argument undergoes macro expansion before substitution, just like an ordinary actual argument: ```text #define A 7 #define V(...) <__VA_ARGS__> #define F(first, ...) first | __VA_ARGS__ V(A, 2, 3) -> <7, 2, 3> F(1, A, 3) -> 1 | 7, 3 ``` The variadic tail may be empty. Both ```text F(1) F(1,) ``` are valid and substitute an empty `__VA_ARGS__`. This does **not** imply that a comma written explicitly in the replacement list is removed automatically. For example, with ```text #define E(format, ...) output(format, __VA_ARGS__) ``` `E("ok")` leaves the comma before the empty tail. The historical GNU `, ## __VA_ARGS__` comma-swallowing behavior is deliberately outside the 0.0.46 contract and remains unsupported; modern code should use `__VA_OPT__(,)`. `__VA_ARGS__` participates in the existing `#` and `##` semantics as a real macro parameter. Stringification uses the raw spelling of the whole variadic tail: ```text #define STRV(...) #__VA_ARGS__ STRV(A, b + c) -> "A, b + c" ``` When adjacent to `##`, the variadic argument is likewise substituted without prescan; the ordinary placemarker, token-concatenation, and rescan rules then apply. For example: ```text #define L(...) pre ## __VA_ARGS__ #define R(...) __VA_ARGS__ ## post L(fix) -> prefix R(fix) -> fixpost ``` If the variadic argument contains multiple preprocessing tokens, only the edge token immediately adjacent to `##` is concatenated and the remaining tokens are preserved, exactly as for an ordinary parameter. An empty variadic tail next to `##` behaves as a placemarker. The name `__VA_ARGS__` is reserved for the variable argument and is not accepted as an ordinary formal parameter name. `#__VA_ARGS__` is valid only in a variadic macro. Dump modes preserve the variadic form of the definition, for example: ```text #define F(first,...) first | __VA_ARGS__ ``` ### 7.4. `__VA_OPT__` Starting with 0.0.47, variadic macros support the standard conditional fragment `__VA_OPT__(pp-tokens)`. If the variable argument contains no preprocessing tokens after normal macro substitution, the entire `__VA_OPT__(...)` expands to an empty sequence. If the variable argument is nonempty, the parenthesized contents participate in the replacement list: ```text #define DEBUG(format, ...) \ fprintf(stderr, format __VA_OPT__(,) __VA_ARGS__) DEBUG("ready") -> fprintf(stderr, "ready") DEBUG("x=%d", x) -> fprintf(stderr, "x=%d", x) ``` Emptiness is decided **after expansion of the variable argument**, not from its raw spelling. Therefore a macro that itself expands to an empty sequence does not activate `__VA_OPT__`: ```text #define EMPTY #define HAS(...) [__VA_OPT__(yes)] HAS() -> [] HAS(EMPTY) -> [] HAS(token) -> [yes] ``` The contents of `__VA_OPT__` may contain balanced nested parentheses. The closing `)` of the `__VA_OPT__` construct is found with nesting taken into account. A nested `__VA_OPT__` inside another `__VA_OPT__` is deliberately forbidden. `__VA_OPT__` is integrated with the existing rules for parameter substitution, stringification, token concatenation, placemarkers, and rescan. For example: ```text #define X 123 #define S(...) #__VA_OPT__(__VA_ARGS__) #define L(...) pre ## __VA_OPT__(__VA_ARGS__) S() -> "" S(X) -> "123" L() -> pre L(X) -> pre123 ``` With `#__VA_OPT__(...)`, parameter substitution inside the fragment happens first, including prescan of ordinary parameters, but arbitrary macro names in the fragment are not additionally rescanned before stringification. Thus: ```text #define X 123 #define S(a, ...) #__VA_OPT__(a X) S(X, y) -> "123 X" ``` If a parameter inside `__VA_OPT__` participates directly in an internal `##`, prescan is suppressed for that parameter in the usual way; the paste is performed before later rescan. An outer `##` adjacent to `__VA_OPT__` receives the edge token of the already prepared fragment. An empty `__VA_OPT__` result next to `##` behaves as a placemarker. `__VA_OPT__` is valid only in the replacement list of a variadic function-like macro and must immediately introduce a parenthesized fragment. `##` cannot be the first or last preprocessing token inside that fragment. The historical GNU extension ```text , ## __VA_ARGS__ ``` is intentionally **not implemented** by `mcpu-cpp`. Use the modern `__VA_OPT__(,)` form for a conditional comma. The old GNU named variadic parameter form `args...` also remains unsupported. ### 7.5. Whitespace normalization in replacement lists Starting with 0.0.48, `mcpu-cpp` does not carry alignment whitespace from a multi-line macro definition into the expansion result. After `\\` + newline has been removed, a whitespace sequence belonging to the replacement list itself is canonicalized to one ASCII space. This is particularly important for definitions whose backslashes are visually aligned in one column: ```text #define TRACE(x) \ do \ { \ output(x); \ done(); \ } \ while( 0 ) ``` Such a definition expands to the compact replacement: ```text do { output(x); done(); } while( 0 ) ``` rather than preserving dozens of spaces before each former physical-line boundary. Normalization applies **only to whitespace belonging to the replacement list**. `mcpu-cpp` is not a source formatter: whitespace in ordinary input text is preserved. Whitespace inside an actual macro argument is likewise not reformatted merely because the argument is substituted into a macro: ```text #define ID(x) x ID(a + b) -> a + b ``` String and character literal contents are preserved verbatim, so: ```text #define S "left right" ``` still contains five spaces inside the string. The presence of whitespace between preprocessing tokens is preserved as one space. This prevents accidental retokenization such as turning `+ +` into `++`, `- >` into `->`, or `< <` into `<<`. The `#` and `##` operators, placemarkers, `__VA_ARGS__`, `__VA_OPT__`, and later rescan keep their existing rules; the policy changes only the amount of ordinary replacement-list whitespace. Dump modes (`-dM`, `-dD`) show the same canonical replacement-list form stored in the internal macro table. ### 7.6. Invisible-line compaction and line markers Starting with 0.0.49, `mcpu-cpp` uses the same model as GNU CPP for vertical whitespace: **remove it, but do not forget it**. Source lines that produce no output preprocessing token after preprocessing need not remain as physical blank lines in the `.E` output, but their source position still contributes to line markers and to `__LINE__`. Why a line is invisible does not matter. It may be a consumed directive, an inactive `#if` branch, a single-line or multi-line comment, an ordinary blank line, or any mixture of these. The emitter compares its current output source position with the position of the next line that will actually be emitted. If the next position is fewer than eight lines away, the gap is represented by ordinary newlines. If the distance is eight lines or greater, the long run of blank lines is replaced by a corrective line marker: ```text # N "file" ``` and the next content line immediately belongs to source line `N`. The behavior therefore matches the GNU CPP boundary: gaps 0 through 7 use newlines; a gap of 8 or more uses a line marker. Structural enter/return markers for included files keep their ordinary meaning: ```text # 1 "header.h" 1 # 4 "source.c" 2 ``` If an included file produces no output, `mcpu-cpp` does not invent a marker reporting how far the preprocessor progressed internally through that header. An enter marker may be followed immediately by its return marker. The real position is corrected again only when some following content must be emitted. This optimization changes only the representation of the output stream. Source coordinates, `__LINE__`, diagnostics, `#line`, include enter/return semantics, and macro processing remain tied to the logical source stream, not to the number of physical lines in the compacted `.E` file. ## 8. Predefined macros Starting with 0.0.6, the historical predefined-macro mechanism was restored in the preprocessor. It is treated as a separate ABI/environment layer for the future unnamed C-like language. These definitions are not decorative: their names and values must match either GNU CPP semantics or an explicitly documented MCPU/LibMPU contract. ### 8.1. Dynamic source macros The following predefined macros are evaluated at the point of use: | Macro | Expansion | |---|---| | `__FILE__` | string constant containing the name of the current input file | | `__LINE__` | decimal number of the current source line | | `__BASE_FILE__` | string constant containing the primary input file name of the translation unit | | `__INCLUDE_LEVEL__` | `#include` nesting level; `0` in the primary file | | `__DATE__` | preprocessor start date in the form `"Mmm dd yyyy"` | | `__TIME__` | preprocessor start time in the form `"hh:mm:ss"` | `__DATE__` and `__TIME__` share one timestamp for the whole translation unit. Their special expansion is emitted without another macro rescan. These names reside in the ordinary macro table, so `#undef` followed by `#define` may deliberately replace a builtin. ### 8.2. Preprocessor version Starting with 0.0.8, the standalone preprocessor does not define GCC's `__VERSION__`. That name belongs to a compiler environment, which does not yet exist for the future high-level language. The version of `mcpu-cpp` has its own unambiguous name: ```text #define __MCPU_CPP_VERSION__ "1.0.3" ``` The value is obtained automatically from `PACKAGE_VERSION`. When a compiler frontend/driver appears, its version contract will be defined separately and will not be mixed with the version of the standalone preprocessor. ### 8.3. ABI sources of truth `mcpu-cpp` is built only with GNU GCC. During `configure`, the project follows the established LibMPU/LibMPUIO `acsite.m4` approach: GCC predefined macros describe native type sizes, byte/word order, and machine-register width, while the installed `` is the final source of truth for LibMPU configuration. In particular, the following values are captured and checked: ```text MPU_REAL_IO_LIMIT MPU_MATH_FN_LIMIT MPU_BYTE_ORDER MPU_WORD_ORDER BITS_PER_MACHINE_REGISTER BITS_PER_UNIT_T sizeof(__mpu_size_t) sizeof(__mpu_ptrdiff_t) ``` `configure` additionally verifies that the byte order and `BITS_PER_MACHINE_REGISTER` recorded by LibMPU agree with the GCC target used to build `mcpu-cpp`. `MPU_WORD_ORDER` is taken directly from the configured LibMPU profile and describes word order in the MCPU data environment. `MPU_REAL_IO_LIMIT` and `MPU_MATH_FN_LIMIT` serve different purposes. For example, a library may support Real I/O up to 65536 bits while providing mathematical functions only up to 16384 bits. Therefore `MPU_MATH_FN_LIMIT` is not used as the limit on existence of Real types. ### 8.4. MCPU architecture and assembler prefixes The target architecture is identified by: ```text #define _ARCH_MCPU 1 ``` MCPU PTR64 is 64 bits wide, so `__SIZEOF_POINTER__`, `__MCPU_POINTER_WIDTH__`, `__INTPTR_TYPE__`, `__UINTPTR_TYPE__`, and the corresponding width/max macros are defined accordingly. Assembler-prefix macros follow GNU CPP meaning rather than the first letter of a register-view name. `mcpu-as` syntax uses no extra sigil before a register, label, or immediate value. The letters `r` and `c` belong to MCPU register syntax; they are not a `REGISTER_PREFIX`. Therefore: ```text #define __REGISTER_PREFIX__ #define __LOCAL_LABEL_PREFIX__ #define __USER_LABEL_PREFIX__ #define __IMMEDIATE_PREFIX__ ``` all four expand to an empty sequence. `.L...` remains a compiler naming convention and is not an assembler-ABI local-label prefix: LOCAL/GLOBAL binding is determined by symbol directives. ### 8.5. Byte order and word order The basic numeric byte-order values are compatible with GNU CPP: ```text __ORDER_LITTLE_ENDIAN__ __ORDER_BIG_ENDIAN__ __ORDER_PDP_ENDIAN__ ``` The target environment publishes its own MCPU names: ```text #define __MCPU_BYTE_ORDER__ __ORDER_LITTLE_ENDIAN__ #define __MCPU_WORD_ORDER__ __ORDER_LITTLE_ENDIAN__ #define __BYTE_ORDER__ __MCPU_BYTE_ORDER__ ``` The actual values of `__MCPU_BYTE_ORDER__` and `__MCPU_WORD_ORDER__` come from the configured LibMPU profile (`MPU_BYTE_ORDER` and `MPU_WORD_ORDER`). They therefore follow the host data representation for which LibMPU was built. This does not alter the separate architectural contract for MCPU instruction bytecode encoding. The GNU/C-specific name `__FLOAT_WORD_ORDER__` is not defined because the future MCPU language has no `float` type. LibMPU/MCPU environment parameters are published in the MCPU namespace: ```text __MCPU_MACHINE_REGISTER_WIDTH__ __MCPU_REAL_IO_LIMIT__ __MCPU_MATH_FN_LIMIT__ __MCPU_INT_MAX_WIDTH__ __MCPU_REAL_MAX_WIDTH__ __MCPU_COMPLEX_MAX_WIDTH__ ``` `__MCPU_INT_MAX_WIDTH__` is `NB_I_MAX * 8`; the Real/Complex maximum width is the configured `MPU_REAL_IO_LIMIT`. `__MCPU_MACHINE_REGISTER_WIDTH__` is the `BITS_PER_MACHINE_REGISTER` value of the installed LibMPU. Real I/O and math limits are deliberately kept separate: `MPU_REAL_IO_LIMIT` controls existence of Real/Complex type families and text conversion, while `MPU_MATH_FN_LIMIT` controls availability of mathematical functions at a given width. ### 8.6. MCPU size/ssize, `ptrdiff`, and pointers The future language does not inherit variable-width C names such as `short`, `int`, and `long`, and it does not use the C-style name `size_t` as part of its own ABI. The unsigned LibMPU size type and signed byte-count/error type are published symmetrically in the MCPU namespace. For a 64-bit configured profile, for example: ```text #define __MCPU_SIZE_TYPE__ uint64 #define __MCPU_SIZE_WIDTH__ 64 #define __MCPU_SIZEOF_SIZE__ 8 #define __MCPU_SIZE_MAX__ 0xffffffffffffffff #define __MCPU_SSIZE_TYPE__ int64 #define __MCPU_SSIZE_WIDTH__ 64 #define __MCPU_SIZEOF_SSIZE__ 8 #define __MCPU_SSIZE_MAX__ 0x7fffffffffffffff ``` This is an MCPU-specific family, not an attempt to invent a nonexistent GNU CPP `__SSIZE_*` contract. The MCPU pointer ABI is independent of the host: PTR64 is always 64 bits wide: ```text #define __INTPTR_TYPE__ int64 #define __UINTPTR_TYPE__ uint64 #define __INTPTR_WIDTH__ 64 #define __UINTPTR_WIDTH__ 64 #define __INTPTR_MAX__ 0x7fffffffffffffff #define __UINTPTR_MAX__ 0xffffffffffffffff #define __SIZEOF_POINTER__ 8 #define __MCPU_POINTER_WIDTH__ 64 ``` The difference between two MCPU pointers is signed and also fixed independently of the host: ```text #define __PTRDIFF_TYPE__ int64 #define __PTRDIFF_WIDTH__ 64 #define __SIZEOF_PTRDIFF__ 8 #define __PTRDIFF_MAX__ 0x7fffffffffffffff ``` Computed MIN expressions such as `(-__PTRDIFF_MAX__ - 1)` are not added to the predefined table. ### 8.7. Character types The future language has no ordinary C `char`. Therefore `__CHAR_TYPE__` and `__WCHAR_TYPE__` are not defined. Language types are named without C/C++ `_t` suffixes: ```text #define __CHAR8_TYPE__ char8 #define __CHAR16_TYPE__ char16 #define __CHAR8_WIDTH__ 8 #define __CHAR16_WIDTH__ 16 #define __SIZEOF_CHAR8__ 1 #define __SIZEOF_CHAR16__ 2 ``` These are types of the future language. The implementation of `mcpu-cpp` itself continues to use LibMPUIO `__mpu_char16_t` and the strict UCS-2 text model internally. ### 8.8. LibMPU integer families Complete structural metadata for integer families is generated up to the actual `NB_I_MAX * 8` of the installed LibMPU rather than stopping at a hard-coded final type. For every power-of-two width starting at 8 bits, TYPE, WIDTH, and SIZEOF are defined: ```text #define __INT1024_TYPE__ int1024 #define __UINT1024_TYPE__ uint1024 #define __INT1024_WIDTH__ 1024 #define __UINT1024_WIDTH__ 1024 #define __SIZEOF_INT1024__ 128 #define __SIZEOF_UINT1024__ 128 ``` With the current LibMPU 1.0.35, `NB_I_MAX == 8192`, so the family extends to `int65536`/`uint65536`, with `__SIZEOF_INT65536__ == 8192`. Decimal-digit metadata is defined for **every** permitted integer width: ```text __INT_DECIMAL_DIG__ __UINT_DECIMAL_DIG__ ``` The value is computed by `mcpu-cpp` integer-only helpers from the known bit width. It is the exact number of decimal digits in the maximum value of the type; neither a sign nor a terminating NUL is included in `DECIMAL_DIG`. For unsigned values the maximum is `2^bits - 1`; for signed values it is `2^(bits-1) - 1`. This differs from LibMPU `_int_digs()`, which estimates a string-buffer size and includes room for a terminating NUL. For example: ```text #define __INT64_DECIMAL_DIG__ 19 #define __UINT64_DECIMAL_DIG__ 20 #define __INT256_DECIMAL_DIG__ 77 #define __UINT256_DECIMAL_DIG__ 78 ``` Only the textual maxima themselves are deliberately limited to widths `bits <= 256`: ```text __INT128_MAX__ __UINT128_MAX__ ``` Maxima are produced through LibMPU `iuitoa()`. Macros named `__INT_MIN__` are not generated: the predefined table must not contain computed expressions such as `(-__INT_MAX__ - 1)`. For widths above 256 bits, only MAX is absent; TYPE/WIDTH/SIZEOF/DECIMAL_DIG continue through the full `NB_I_MAX * 8` range. ### 8.9. LibMPU Real and Complex families Real/Complex structural metadata is generated for every power-of-two width from 32 bits through the actual configured `MPU_REAL_IO_LIMIT`. TYPE, WIDTH, and SIZEOF are published for all of these types. For Complex, WIDTH denotes the type parameter, not total storage width: ```text #define __COMPLEX128_TYPE__ complex128 #define __COMPLEX128_WIDTH__ 128 #define __SIZEOF_COMPLEX128__ 32 ``` `complex128` consists of two `real128` components, so its storage size is 32 bytes. With `MPU_REAL_IO_LIMIT == 65536`, the top of the family is: ```text #define __COMPLEX65536_TYPE__ complex65536 #define __COMPLEX65536_WIDTH__ 65536 #define __SIZEOF_COMPLEX65536__ 16384 ``` For Real: ```text #define __REAL65536_TYPE__ real65536 #define __REAL65536_WIDTH__ 65536 #define __SIZEOF_REAL65536__ 8192 ``` Precision metadata is defined for **all** allowed Real widths up to `MPU_REAL_IO_LIMIT`. Macro names correspond directly to LibMPU helpers: ```text __REAL_DECIMAL_DIG__ -> _real_digs(bits/8) __REAL_MANT_DIG__ -> _real_mant_digs(bits/8) ``` `__REAL_DIG__` is intentionally absent. The `bits <= 256` restriction applies only to large textual numeric constants. For widths up to 256 bits, the following are also defined: ```text __REAL_MAX__ __REAL_MIN__ __REAL_EPSILON__ __REAL_MAX_EXP__ __REAL_MIN_EXP__ __REAL_MAX_10_EXP__ __REAL_MIN_10_EXP__ ``` For example, with LibMPU 1.0.35, the current `real128` profile gives values of the form: ```text #define __REAL128_EPSILON__ 2.524354896707237777317531409e-29 #define __REAL128_MAX__ 4.197157432934775384808581951e+323228496 #define __REAL128_MIN__ 9.530259619551804292864984035e-323228497 #define __REAL128_MAX_10_EXP__ 323228496 #define __REAL128_MAX_EXP__ 1073741823 #define __REAL128_MIN_10_EXP__ -323228524 #define __REAL128_MIN_EXP__ -1073741822 ``` MAX/MIN/EPSILON are created by LibMPU itself and converted through `real_to_ascii()`. Exponent constants are obtained from LibMPU exponent helpers and integer conversion. For widths above 256 bits, these numeric predefines are absent, but TYPE/WIDTH/SIZEOF/DECIMAL_DIG/MANT_DIG continue through `MPU_REAL_IO_LIMIT`. For every supported Real type through `MPU_REAL_IO_LIMIT`, two compact characteristics are also published: ```text #define __SIZEOF_REAL128_EXP__ 4 #define __REAL128_MAX_STRLEN__ 60 ``` `__SIZEOF_REALxxx_EXP__` is obtained directly from `_sizeof_exp(NB_Rxxx)`. `__REALxxx_MAX_STRLEN__` comes from `_real_max_string(NB_Rxxx)` and is the maximum **number of characters** in the textual representation, not a byte count. A zero-terminated string therefore needs at least `__REALxxx_MAX_STRLEN__ + 1` elements: for `char8` that is the same number of bytes, while for `char16` the physical byte count is twice as large. These two metadata macros are also defined for Real widths above 256 bits because their own values remain small. ### 8.10. Macro dumps: `-dM`, `-dMP` The command: ```text mcpu-cpp -dM input.c ``` prints only final **non-predefined** macros in `#define ...` form. This group includes definitions from the primary file and included headers, as well as command-line `-D` definitions. MCPU-CPP's own predefined macros are not printed by `-dM`. This mode is therefore intended primarily for a compact inspection of macro state created by the user program. The command: ```text mcpu-cpp -dMP input.c ``` adds active MCPU-CPP predefined macros to the same final state. Output contains two consecutive groups: predefined macros first, then non-predefined macros. Definitions inside each group are sorted deterministically by name. This is useful for system development because it exposes the preprocessing ABI and architectural properties of the current MCPU environment without mixing them with user definitions. Group membership is determined by macro origin, not by spelling. A macro created by `-D` or `#define` is ordinary even if its name looks system-like. If a predefined macro is removed with `#undef`, it is not printed. If the user then defines the same name again, the new definition belongs to the ordinary group and appears in the corresponding part of `-dMP`, and also in `-dM`. Thus both modes display the **final macro state**. Context-dependent `__FILE__`, `__LINE__`, `__DATE__`, `__TIME__`, `__BASE_FILE__`, and `__INCLUDE_LEVEL__` are not printed by the static dump. Static ABI/architecture predefined macros and computed static Real metadata are printed by `-dMP`. When an input file is supplied, it is fully preprocessed first and the final macro state is printed afterward; ordinary preprocessed text is not emitted in `-dM`/`-dMP` modes. Without an input file, stdin is used, so empty stdin with `-dM` gives an empty dump while `-dMP` provides the active static predefined macros of the current MCPU environment. `-dD` has different semantics and is unaffected by this distinction. ### 8.11. Definition dump: `-dD` The command: ```text mcpu-cpp -dD input.c ``` preserves ordinary preprocessing output and additionally emits encountered `#define` directives. Before primary input starts, static predefined macro definitions are printed. Each such definition is preceded by a marker: ```text # 0 "" #define NAME value ``` and the predefined block itself is preceded by an input-file marker of the form `# 0 "input.c"`. Context-dependent `__FILE__`, `__LINE__`, `__DATE__`, `__TIME__`, `__BASE_FILE__`, and `__INCLUDE_LEVEL__` are not included in the initial built-in block. ### 8.12. Configuration dump: `-dconfig` The command: ```text mcpu-cpp -dconfig ``` requires no input file and prints the effective configuration-variable layer after runtime configuration, optional system override, home user override, or a selected `--config-file` have been read, including `$NAME`/`${NAME}` expansion. With `--sys-root=PATH`, configuration files are not read and the dump instead exposes the command-line-selected `PATH/include` system root. Lines are sorted by name and printed as: ```text NAME = value; ``` This makes it possible to inspect actual include paths without manually searching `/etc/mcpu-cpp.conf`, `/etc/mcpu/mcpu-cpp.conf`, and `$HOME/.mcpu/etc/mcpu-cpp.conf`. ### 8.13. Verbose configuration snapshot: `-v` With `-v`, MCPU-CPP retains its runtime trace for `#lang`, `#include`, and `#include_next`, but configuration variables are printed only once, after all configuration layers have been read and priority rules applied. Verbose output therefore shows only **effective values**; intermediate values from runtime root, system, and user configuration are not duplicated. The configuration block follows include-policy order: language-specific user paths, the common user path, the system root, and the AFTER path. A variable that is absent from every configuration layer is not printed. The runtime-derived default `MCPU_CPP_SYSTEM_INCLUDE_PATH` is a full lowest-priority value and is therefore visible under `-v` even when no `mcpu-cpp.conf` exists **or all configuration files are disabled with `--no-config`**. The line form is: ```text config: NAME=value ``` ### 8.14. Effective search directories: `-dsearch-dirs` The command: ```text mcpu-cpp -dsearch-dirs ``` requires no input file, prints the effective global search directories, and exits without preprocessing. The format is intentionally simple: ```text search: /path/to/directory ``` Directories are printed in semantic search-class order: ```text explicit -I explicit -isystem configured language-specific user directories MCPU_CPP_INCLUDE_PATH MCPU_CPP_SYSTEM_INCLUDE_PATH/ MCPU_CPP_SYSTEM_INCLUDE_PATH explicit -idirafter MCPU_CPP_AFTER_INCLUDE_PATH ``` Language-specific entries are printed for every supported language in their canonical order. During a real `#include`, only the directory corresponding to the active `#lang` participates. The directory of the current physical file is not printed by `-dsearch-dirs`: it exists only dynamically for a particular `#include "..."` and changes with the include stack. `--no-config` does not remove the runtime-derived system root, so even without configuration files the dump still contains `/include/` and `/include`. `-nostdinc` removes the effective system `` entries and system root from the dump, but not explicit `-isystem`. A directory that does not exist in the file system is still displayed because it remains part of the effective search configuration and will simply be skipped during a real file search. `-dsearch-dirs` accounts for `-I`, `-isystem`, `-idirafter`, every configuration layer, and the replacement semantics of `MCPU_CPP_SYSTEM_INCLUDE_PATH`. Using `-o` with this action is an error. ### 8.15. Conditional compilation The directives `#if`, `#ifdef`, `#ifndef`, `#elif`, `#else`, and `#endif` are processed as preprocessor control directives and are never copied to the output stream, including under `-dD`. Inactive branches are skipped without executing `#define`, `#undef`, or `#include` directives within them; nested conditional groups are still tracked correctly. An `#if` expression first processes the `defined` operator, then undergoes macro expansion, and any remaining identifiers evaluate to `0`. Arithmetic, bitwise, comparison, and logical operators are supported, as are `?:` and short-circuit semantics for `&&`, `||`, and `?:`. Starting with 0.0.26, expression syntax is parsed by a parser generated by ZUBR 4.1.0 from `src/mcpp-expr.zubr`; the same file contains the UCS-2 lexical analyzer. `defined` preprocessing and macro expansion take place before the parser is entered. Arithmetic semantics live in `mcpp-semantic.c/h` and do not depend on the integer sizes of the host system. Generated `mcpp-expr.c` is included in releases, so ZUBR is required only when the grammar changes. #### 8.15.1. The only evaluation width is 64 bits MCPU-CPP is a preprocessor, not a general-purpose language compiler. All integer computation in conditional directives uses only 64-bit arithmetic. The preprocessor does not perform arbitrary-width LibMPU arithmetic, floating point, or complex-number computation. When a programmer does not need explicit control over the binary representation of a literal, ordinary integer constants with optional `U`/`u` are sufficient. For example: ```c #if 2 > 1 #if 0xffffffffffffffffU > 1 ``` A numeric lexeme remains in UCS-2 until classification, after which its ASCII portion is passed to LibMPU `iatoui()`. Binary `0b...`, octal `0...`, decimal, and hexadecimal `0x...` forms are supported. A value that does not fit in 64 bits is an error. Old C suffixes `L`, `l`, `LL`, and `ll` are not supported. #### 8.15.2. Width suffix `zNNN[Uu]` MCPU-CPP understands the width suffix shared by MCPU languages: ```text zNNN ZNNN zNNNu zNNNU ZNNNu ZNNNU ``` `NNN` is a nonempty sequence of decimal digits and is **always** interpreted in decimal, even with leading zeroes. Thus `z8`, `z08`, and `z008` all denote the same width of 8 bits. In the general MCPU syntax, a valid width must be a power of two from 8 through `MPU_REAL_IO_LIMIT`. MCPU-CPP, however, deliberately limits evaluation to 64 bits: * `z8`, `z16`, `z32`, `z64`, in either letter case, are valid; * `NNN > 64` is immediately an error: conditional preprocessing does not accept numeric constants wider than 64 bits; * if `NNN <= 64` but is not a valid power-of-two width, such as `z24`, a warning is issued and the `zNNN` part itself is ignored; * a following optional `U`/`u` selects unsigned interpretation and retains that meaning even when an invalid `zNNN` has been ignored. The numeric preprocessing token must end after the complete suffix. An operator or punctuation character begins the next token, so `1z32u+2`, `(1z32u)`, and `1z32u==1` are valid. Forms such as `1z32undefined`, `1z32ufoo`, and `1z32$foo` are errors and are not artificially split into a number followed by a name. #### 8.15.3. Literal normalization The width suffix acts **exactly once, while the value of the literal itself is formed**. The width is not retained in the semantic value and has no role in later operations. For `VALUEzNNN`, the value is treated as a signed N-bit two's-complement number: 1. retain the low `NNN` bits; 2. sign-extend the result to 64 bits. For `VALUEzNNNu`/`VALUEzNNNU`, the low `NNN` bits are retained and then zero-extended to 64 bits. For example: ```text 0x7fz8 -> 0x000000000000007f -> 127 0x80z8 -> 0xffffffffffffff80 -> -128 0xffz8 -> 0xffffffffffffffff -> -1 0x80z8u -> 0x0000000000000080 -> 128 0xffz8u -> 0x00000000000000ff -> 255 0x1ffz8 -> 0xffffffffffffffff -> -1 0x1ffz8u -> 0x00000000000000ff -> 255 ``` The last two examples are deliberate: `zNNN` specifies the width of the **binary representation**, not a mathematical range check. Bits above N are discarded before extension. After this normalization there is no remaining `z8`, `z16`, or `z32` concept in the evaluation model. The internal value contains only a 64-bit bit pattern and signed/unsigned state. #### 8.15.4. All subsequent operations are 64-bit After literal normalization, every arithmetic, bitwise, comparison, and logical operation uses 64-bit operands. An operation result is not truncated back to the width of the original suffix. Therefore: ```text 0x7fz8 + 1 -> 128 0xffz8u + 1 -> 256 ``` not `-128` and `0`. Likewise `~0xffz8u` inverts all 64 bits and gives `0xffffffffffffff00`. For binary operations where signedness matters, the presence of an unsigned operand selects 64-bit unsigned interpretation. Comparisons return `0` or `1`. Logical `!`, `&&`, and `||` also return signed 64-bit `0` or `1`; short-circuit evaluation does not evaluate an unselected operand. Shifts happen after 64-bit normalization. Right shift of a negative signed value is arithmetic; right shift of an unsigned value is logical. For example: ```text 0x80z8 >> 1 -> -64 0x80z8u >> 1 -> 64 ``` The historical MCPU-CPP rule for a negative shift count is preserved: `A << -N` is equivalent to `A >> N`, and `A >> -N` is equivalent to `A << N`. Thus `zNNN` does not turn the preprocessor into a compiler with integer promotions over multiple widths. It only allows the binary representation of the source literal to be stated explicitly; the expression then evaluates in one simple 64-bit model. #### 8.15.5. Character constants A character unit has type `__mpu_uint16_t`, matching the internal UCS-2 representation, and is zero-extended to 64 bits before evaluation. Subsequent arithmetic is again ordinary 64-bit arithmetic. Conditional-compilation state is stored on a separate stack; a conditional group may not cross an include-file boundary. ### 8.16. Diagnostic directives `#error` and `#warning` MCPU-CPP supports the standard diagnostic directives: ```text #error message #warning message ``` `#error` emits an error diagnostic using the current logical file name and line number and immediately terminates preprocessing unsuccessfully. `#warning` emits a warning with the same source-location information and preprocessing continues. A preceding `#line` therefore affects both diagnostics. The remainder of the line after the directive name **does not undergo macro expansion**. For example: ```c #define MESSAGE expanded #warning MESSAGE ``` prints `MESSAGE`, not `expanded`. This distinguishes diagnostic directives from `#if` and `#line`, where macro expansion is part of the relevant contract. Comments are removed by the ordinary preprocessing phase before the directive is processed. Leading and trailing whitespace in the message is removed and whitespace sequences between preprocessing tokens are collapsed to one space. Whitespace inside quotes is preserved. For example: ```c #warning one /* comment */ two #warning "a b" ``` produce `one two` and `"a b"`, respectively. Unicode text passes through the internal UCS-2 representation and is written to the external diagnostic as UTF-8. Both directives are control directives and are never copied to normal output or to `-dD`. In an inactive `#if` branch they are ignored completely, so the usual protective pattern behaves as expected: ```c #if 0 #error this error is inactive #endif ``` ### 8.17. Warning control: `-Wcomment`, `-Wall`, `-Werror` MCPU-CPP distinguishes mandatory warnings that are part of established preprocessing semantics from optional warning classes enabled by the user. Warning control does not alter `-dD`, macro expansion, conditional compilation, or include search semantics. `-Wcomment` and `-Wcomments` are exact aliases and enable two lexical warnings: * a `/*` sequence seen while already inside an open `/* ... */` comment; * backslash-newline inside a `//` comment, causing that single-line comment to continue physically onto the next source line. This optional class is disabled by default. `-Wall` enables all optional MCPU-CPP warning classes; in version 0.0.40 this class is `-Wcomment`. `-Wno-comment` and `-Wno-comments` disable it. As in the GNU warning model, a more specific setting has priority over a group setting regardless of argument order. Therefore both: ```text mcpu-cpp -Wall -Wno-comment file.c mcpu-cpp -Wno-comment -Wall file.c ``` leave comment warnings disabled. Between settings of equal specificity, the last option wins; for example `-Wno-comment -Wcomment` enables the class. `-Werror` does not enable any new warning class. It promotes to an error every warning that would actually be emitted during that invocation, causing an unsuccessful result. This applies both to optional comment warnings and to existing mandatory MCPU-CPP warnings, including: * an active `#warning` directive; * an invalid `zNNN` width not exceeding 64 bits; * redefinition of a macro with a different replacement list; * a `##` result that does not form a single preprocessing token. For example: ```text mcpu-cpp -Wcomment -Werror file.c ``` turns a detected comment warning into an error. By contrast, `-Werror` alone, without `-Wcomment`/`-Wall`, does not cause MCPU-CPP to search for optional comment warnings. `-Wno-error` restores ordinary warning severity. Between `-Werror` and `-Wno-error`, which have the same specificity, the last command-line option wins. Thus `-Werror -Wno-error` leaves warnings as warnings, while `-Wno-error -Werror` promotes them again. Version 0.0.40 deliberately did not introduce `-Werror=`, `-Wno-error=`, `-Wundef`, `-Wunused-macros`, `-Wtraditional`, or other compiler-oriented classes. The MCPU-CPP warning interface remains compact and is extended only when a class is actually needed by the preprocessing language itself. ### 8.18. UCS-2 identifiers Starting with 0.0.22, preprocessing identifiers are no longer restricted to ASCII. Inside `mcpu-cpp`, text is already strict UCS-2, and characters are classified by locale-independent LibMPUIO 1.0.4 functions based on Unicode 18.0.0. The first identifier character must be `_` or have the `XID_Start` property; following characters must be `_`, `$`, or have `XID_Continue`. `$` is an `mcpu-cpp` extension: it is allowed only after the first character and may not start an identifier. One rule is used consistently for macro names and parameters, `#undef`, `#ifdef`/`#ifndef`, `defined`, ordinary macro expansion, and `#`/`##`. Names remain case-sensitive. Surrogate code units `U+D800..U+DFFF` are not valid identifier characters. For example, all of these are valid: ```c #define АНДРЕЙ 1 #define résumé 2 #define ΩМЕГА 3 #define VALUE$OLD 4 ``` `VALUE$OLD` is valid, while `$VALUE` is invalid because `$` is not an identifier-start character. Combining marks and non-ASCII decimal digits may appear in `XID_Continue` positions but do not automatically become valid initial characters. Numeric constant syntax is unaffected: it follows the rules of the active language, not Unicode `isdigit`. ### 8.19. Command-line macros `-D` and `-U` Starting with 0.0.23, `-D` and `-U` are full preprocessor actions. Supported forms are: ```text -DNAME -DNAME=VALUE -D'FUNC(a,b)=a+b' -UNAME ``` `-DNAME` is equivalent to `#define NAME 1`; an `=` with an empty right-hand side defines an empty replacement list. Function-like command-line definitions use the same macro engine as ordinary `#define`, including parameters, `#`, `##`, and subsequent rescanning. `-U` uses the same identifier contract as `#undef`. `-D`/`-U` actions are executed in command-line order after predefined macros have been installed. Only the payload of `-D` and `-U` is interpreted as UTF-8 and converted to strict UCS-2. File names, `-I`, other pathname arguments, and all other command-line arguments remain the original byte strings and undergo no Unicode conversion. Starting with 0.0.25, `$` is allowed inside a macro name, but not in its first position. When `$` is passed through a shell, the user must account for shell rules: the shell processes `$` **before `mcpu-cpp` starts**. Single quotes fully protect `$`, for example: ```sh mcpu-cpp '-DАНДРЕЙ$_Y=62' input.c ``` Without quotes, `$` must be escaped: ```sh mcpu-cpp -DАНДРЕЙ\$_Y=62 input.c ``` or double quotes may be used with escaping: ```sh mcpu-cpp -D"АНДРЕЙ\$_Y=62" input.c ``` The unprotected form: ```sh mcpu-cpp -DАНДРЕЙ$_Y=62 input.c ``` does not pass the spelling literally: `$...` is expanded by the shell first, and `mcpu-cpp` receives the already modified `argv`. Inside single quotes, a backslash before `$` is unnecessary and would become an ordinary argument character. For command-line `-D`, the left-hand side up to the first `=` is parsed as a separate macro declarator. If an invalid tail occurs after a valid name (or a completed formal-parameter list of a function-like macro) but before `=`, that tail is silently discarded and **never becomes part of the replacement list**. For example: ```text -D'АНДРЕЙ@XYZ=62' ``` is equivalent to: ```c #define АНДРЕЙ 62 ``` not the invalid `#define АНДРЕЙ @XYZ 62`. The same valid-identifier-prefix rule applies to `-U`. If the very first character is not a valid identifier-start character, such as `$` or a digit, the definition remains an error. Under `-dD`, definitions originating from `-D` are marked separately from predefined macros: ```text # 0 "" #define NAME value ``` while predefined macros continue to use ``. ### 8.20. Public command-line interface `mcpu-cpp` supports only the current options documented by `--help`. Obsolete compatibility flags do not form a hidden interface and are diagnosed as `unknown option`. `-E` is the exception: it is silently accepted and ignored because a compiler driver may pass it while invoking a standalone preprocessor. `--object-suffix SUFFIX` selects the object-target suffix used when generating Make dependencies; the argument is mandatory. ## 9. Build system and generators Project-owned Autoconf macros live in the root `acsite.m4`. The `m4/` directory is reserved for external/vendor M4 files. This follows the convention used by MCPU libraries and keeps project configure code separate from imported macros. The `#if` expression parser is generated by ZUBR 4.1.0 from `src/mcpp-expr.zubr`. Release archives contain both the grammar and the already generated `src/mcpp-expr.c`, so an ordinary release build does not require ZUBR. After the grammar changes, a developer build uses the normal Automake rule: ```text zubr -vl -s -Bmcpp_ -o mcpp-expr.c mcpp-expr.zubr ``` Before a release, generated C must correspond to the grammar, and the full test suite plus `make distcheck` must complete without errors. ### 9.1. Developer bootstrap and the Git source tree Starting with 0.0.50, the root `./bootstrap` script makes it unnecessary to store files in Git when they are completely reproducible from source. The script first generates `src/mcpp-expr.c` from `src/mcpp-expr.zubr` using ZUBR 4.1.0, then runs `aclocal`, `autoheader`, `automake`, and `autoconf` in the style used by LibMPU and LibMPUIO. `--target-dest-dir=DIR` selects a target ROOTFS used for the system Autoconf macro/include directories. This convention applies specifically to the developer Git tree. **Release archives remain self-contained**, exactly as before: they contain `configure`, `Makefile.in`, Automake helper scripts, and the generated `src/mcpp-expr.c`. Therefore an ordinary release build requires neither a preliminary `bootstrap` run nor ZUBR. The root `.gitignore` lists reproducible bootstrap files and ordinary configure/build state. It does not change the existing release/build model; it only allows a cleaner Git repository. ## 10. GNU-compatible features `mcpu-cpp` is an independent MCPU preprocessor, but it intentionally follows GNU CPP behavior for a number of well-known operations. Compatibility applies to the documented features; it does not imply complete CLI or language interchangeability with GCC. GNU-compatible behavior is used for, in particular: * object-like and function-like macros, macro rescan, `#`, and `##`; * variadic macros `...` / `__VA_ARGS__` and standard `__VA_OPT__`; * `#if`, `#ifdef`, `#ifndef`, `#elif`, `#else`, `#endif`, and `defined`; * `#include`, `#include_next`, `#pragma once`, `#line`, and GNU linemarkers; * compact output mapping: up to seven invisible lines are represented by newlines, while a gap of eight or more uses a corrective linemarker; * forced files `-include` / `-imacros` and dependency options `-M`, `-MM`, `-MD`, `-MMD`, `-MF`, `-MT`, `-MQ`, and `-MG`; * warning controls `-w`, `-Wall`, `-Werror`, and the supported `-Wcomment` forms. MCPU-specific facilities, including `#lang` / `#endlang`, the `zNNN` numeric suffix, and ABI predefined macros, remain native `mcpu-cpp` extensions.