diff options
| author | kx <kx@radix-linux.su> | 2026-10-01 12:02:28 +0300 |
|---|---|---|
| committer | kx <kx@radix-linux.su> | 2026-10-01 12:02:28 +0300 |
| commit | ba1b04d64bdfaf377915b22f77520216ebe47674 (patch) | |
| tree | 044343035c28b2b04ca9725f26e9b7c39342d19c /doc | |
| parent | 124b140798456778e2a96851cc7268aa9a726698 (diff) | |
| download | mcpu-cpp-1.0.2.tar.xz | |
Version 1.0.21.0.2
Diffstat (limited to 'doc')
| -rw-r--r-- | doc/Makefile.am | 3 | ||||
| -rw-r--r-- | doc/mcpu-cpp-en.md | 2169 | ||||
| -rw-r--r-- | doc/mcpu-cpp-ru.md | 2167 |
3 files changed, 4339 insertions, 0 deletions
diff --git a/doc/Makefile.am b/doc/Makefile.am new file mode 100644 index 0000000..7f354f3 --- /dev/null +++ b/doc/Makefile.am @@ -0,0 +1,3 @@ +EXTRA_DIST = \ + mcpu-cpp-ru.md \ + mcpu-cpp-en.md diff --git a/doc/mcpu-cpp-en.md b/doc/mcpu-cpp-en.md new file mode 100644 index 0000000..5429456 --- /dev/null +++ b/doc/mcpu-cpp-en.md @@ -0,0 +1,2169 @@ +# mcpu-cpp + +`mcpu-cpp` is the preprocessor for MCPU programming languages. It is an +independent component of the LibMPU/LibMPUIO/LibMCPU ecosystem and is not tied +to the name of any single language: the active language is selected with the +`#lang` directive. + +This document defines the normative behavior of `mcpu-cpp`: its text model, +directives, macro engine, include pipeline, configuration, diagnostics, and +dependency generation for MCPU tools. + +## 1. Text model + +External source files and configuration files are encoded in UTF-8. The UTF-8 +must be valid. For source programs, the check that characters belong to the +UCS-2 range is performed after comments have been removed. Therefore a valid +Unicode scalar value above `U+FFFF` is permitted inside a comment, but remains +an error in program text. After this stage, source text is processed as a +sequence of `__mpu_char16_t` values. An input UTF-8 BOM is accepted and +removed. An embedded NUL in a source file is forbidden. + +`CRLF` and `CR` line endings are normalized to `LF`. + +## 2. Actions performed independently of directives + +`mcpu-cpp` performs several transformations before directives are parsed. + +### 2.1. Backslash-newline + +A `\\` immediately followed by a newline is removed before comments, +directives, and macros are recognized. For example, + +```text +#defi\ +ne FOO 10\ +20 +``` + +is equivalent to the logical line + +```text +#define FOO 1020 +``` + +Physical line numbers continue to contribute to the current source position. +Unless the user changes that position with `#line`, those physical positions +are the ones reflected in generated line markers. + +### 2.2. Comments + +`/* ... */` and `// ...` comments are removed before subsequent processing. +Where needed to keep adjacent tokens separate, a whitespace separator is +preserved. If a comment terminates a nonempty line, neither a synthetic +separator nor whitespace that immediately preceded the comment is retained +after the comment is removed: the line ends at its last significant +character. The same rule applies to a multi-line comment that starts after +program text. If comment removal leaves a line containing only whitespace, the +line becomes genuinely empty. A comment between two tokens still leaves the +separator required to keep the tokens from being joined. Newlines are +preserved so source coordinates are not destroyed. + +Comments are not recognized inside string or character constants. In the +`diff` language, an apostrophe is not treated as the beginning of a character +constant because it is used in derivative notation. + +Within a literal `#include <...>` operand, `/*` and `//` sequences are treated +as part of the file name. + +## 3. Directives and the output stream + +A directive begins with `#` when only whitespace or comments precede it on the +logical line. Whitespace is permitted between `#` and the directive name. + +Source-position service information in the output stream uses GNU **line +markers**: + +```text +# line-number "file-name" [flags] +``` + +This is not the input directive `#line`. Entering an included file adds flag +`1` to the line marker, and returning to the file that contained the +`#include` adds flag `2`. These values have the same meaning as in GNU CPP: +`1` means entering a new file and `2` means returning to the previous file. +Flag `2` is not a nesting count or include level. + +For example: + +```text +# 1 "main.c" +# 1 "defs.h" 1 +... +# 2 "main.c" 2 +``` + +The input directive + +```text +#line 62 "main.y" +``` + +is not copied to the output stream. It changes the logical values of +`__LINE__` and `__FILE__` for subsequent text and is represented in output by +a line marker: + +```text +# 62 "main.y" +``` + +The arguments of `#line` undergo macro expansion according to the line-control +model. If an `#include` follows such a `#line`, the return marker receives flag +`2`, for example `# 65 "main.y" 2`. A name installed by `#line` becomes the +logical name used by `__FILE__` and line markers; it does not change the +directory used to resolve a quoted `#include`. + +Preprocessor directives use canonical English names only. Unicode remains fully supported in identifiers, strings, comments, and other user text. + +## 4. Header files + +The following forms are supported: + +```text +#include "file" +#include <file> +#include_next "file" +#include_next <file> +#pragma once +``` + +For ordinary `#include "file"`, the directory of the **physical** current +source file is always checked first. A logical name established by `#line` +does not affect this step. For `#include <file>`, the directory of the current +file is not checked. + +### 4.1. Relocatable MCPU root as an ecosystem-wide principle + +Starting with release 0.0.37, the MCPU installation directory **does not +contain a version number of a particular tool** and is not an absolute runtime +constant compiled into the binary. A version belongs to `mcpu-cpp`, +`mcpu-as`, `mcpu-ld`, `mcpu-run`, or a library; it does not define the root of +the shared MCPU environment. + +For a typical configuration: + +```text +./configure --prefix=/usr --libdir=/usr/lib64 +``` + +`make install` creates: + +```text +/usr/lib64/mcpu/ +├── bin/ +│ └── mcpu-cpp +├── etc/ +│ └── mcpu-cpp.conf +├── include/ +│ ├── diff/ +│ ├── dift/ +│ ├── alg/ +│ ├── as/ +│ ├── avm/ +│ └── acs/ +└── lib/ # common directory for future MCPU libraries +``` + +The public program name lives in `$bindir`: + +```text +/usr/bin/mcpu-cpp -> ../lib64/mcpu/bin/mcpu-cpp +``` + +The absolute `/usr/lib64/mcpu` path is **not part of the MCPU-CPP runtime +ABI**. It is only the configure-time installation location selected by +`make install`. + +On every normal invocation, MCPU-CPP determines the actual path of its own +executable through Linux `/proc/self/exe`. The public-command symlink does not +interfere with this: `/proc/self/exe` names the binary that is actually being +executed. If `/proc/self/exe` is unavailable, a fallback resolves `argv[0]` +through `PATH` and `realpath(3)`; there is no fallback to a compiled-in +configure-time installation root. + +For an executable + +```text +<root>/bin/mcpu-cpp +``` + +the runtime root is derived as: + +```text +executable = <root>/bin/mcpu-cpp +executable dir = <root>/bin +MCPU runtime root = <root> +``` + +and the following paths are derived from it automatically: + +```text +<root>/etc/mcpu-cpp.conf +<root>/include +``` + +Therefore the whole tree can be physically moved, for example from + +```text +/usr/lib64/mcpu/ +``` + +to + +```text +/opt/mcpu-test/ +``` + +or + +```text +$HOME/devel/mcpu-next/ +``` + +and `<new-root>/bin/mcpu-cpp` immediately starts using +`<new-root>/etc/mcpu-cpp.conf` and `<new-root>/include` without being +reconfigured. The old absolute path is retained neither in runtime defaults +nor in the installed `mcpu-cpp.conf`. + +This is not a preprocessor-specific trick; it is a **general MCPU ecosystem +principle**. Future `mcpu-as`, `mcpu-ld`, `mcpu-run`, libraries, CRT, and other +components are expected to share one relocatable root: + +```text +<root>/bin +<root>/etc +<root>/include +<root>/lib +``` + +Their own versions may differ. Consistency of a particular MCPU environment is +defined by all components residing in one runtime tree, not by matching +version suffixes in directory names. + +### 4.2. Runtime defaults, configuration layers, and the system include root + +Before reading any configuration file, MCPU-CPP creates the runtime-derived +value: + +```text +MCPU_CPP_SYSTEM_INCLUDE_PATH = <runtime-root>/include +``` + +Configuration layers are then applied in order of increasing priority: + +```text +runtime-derived defaults + ↓ +<runtime-root>/etc/mcpu-cpp.conf + ↓ +/etc/mcpu/mcpu-cpp.conf + ↓ +$HOME/.mcpu/mcpu-cpp.conf +``` + +`<runtime-root>/etc/mcpu-cpp.conf` is installed with MCPU-CPP, but deliberately +does not contain an absolute default `MCPU_CPP_SYSTEM_INCLUDE_PATH`: otherwise +moving the tree would restore the old path. `/etc/mcpu/mcpu-cpp.conf` is an +optional machine-wide override; `make install` does not create `/etc/mcpu`. +`$HOME/.mcpu/mcpu-cpp.conf` is also optional, is not versioned, and has the +highest configuration priority. + +If one variable is defined more than once, the last definition wins, including +an empty definition. Therefore `MCPU_CPP_SYSTEM_INCLUDE_PATH` remains a fully +replaceable system root. For example: + +```text +MCPU_CPP_SYSTEM_INCLUDE_PATH = $HOME/mcpu-next/include; +``` + +completely replaces the runtime-derived `<runtime-root>/include`. For active +`#lang "as"`, the following locations are then checked: + +```text +$HOME/mcpu-next/include/as +$HOME/mcpu-next/include +``` + +Standard language subdirectories are always derived by the preprocessor from +one root; there are no variables named +`MCPU_CPP_SYSTEM_<LANG>_INCLUDE_PATH`. + +An empty effective value: + +```text +MCPU_CPP_SYSTEM_INCLUDE_PATH = ; +``` + +removes the configured system stage entirely. A higher-priority configuration +file may later enable it again with a nonempty value. + +`--config-file FILE` applies an explicitly selected file on top of the +runtime-derived default. `--no-config` disables **only configuration-file +reading**: `<runtime-root>/etc/mcpu-cpp.conf`, `/etc/mcpu/mcpu-cpp.conf`, and +`$HOME/.mcpu/mcpu-cpp.conf` are not read, but `<runtime-root>/include` remains +the standard system root. Only `-nostdinc` removes the effective standard +system tree from include search for one invocation; an explicit `-isystem` +still remains a command-line directory. + +### 4.3. Normative include-file search order + +Search order is part of the MCPU-CPP contract. Explicit command-line +parameters have priority over persistent configuration. After the optional +directory of the current physical file, the effective chain is strictly: + +```text +explicit -I + ↓ +explicit -isystem + ↓ +MCPU_CPP_<LANG>_INCLUDE_PATH + ↓ +MCPU_CPP_INCLUDE_PATH + ↓ +MCPU_CPP_SYSTEM_INCLUDE_PATH/<lang> + ↓ +MCPU_CPP_SYSTEM_INCLUDE_PATH + ↓ +explicit -idirafter + ↓ +MCPU_CPP_AFTER_INCLUDE_PATH +``` + +Entries that are absent or do not contain the requested file are skipped. + +`MCPU_CPP_<LANG>_INCLUDE_PATH` denotes user-configurable language-specific path +lists: + +```text +MCPU_CPP_DIFF_INCLUDE_PATH +MCPU_CPP_DIFT_INCLUDE_PATH +MCPU_CPP_ALG_INCLUDE_PATH +MCPU_CPP_AS_INCLUDE_PATH +MCPU_CPP_AVM_INCLUDE_PATH +MCPU_CPP_ACS_INCLUDE_PATH +``` + +The user fully controls the names and locations of these directories. +`MCPU_CPP_INCLUDE_PATH` is a common user path list visible in every language +state. + +`-idirafter` and `MCPU_CPP_AFTER_INCLUDE_PATH` form a common fallback area. +MCPU-CPP does not automatically derive `<lang>` subdirectories for them. The +user controls their internal layout and may, for example, write: + +```text +#include <vendor/device.h> +``` + +Priority is determined by the semantic class, not by the relative appearance +of different classes in argv or configuration. Within one class, insertion +order is preserved. + +### 4.4. `#include_next` and wrapper headers + +`#include_next` is intended primarily for wrapper headers. It allows a local +header to precede a system header, adjust local policy, and then continue the +search for a same-named header along the normative chain without copying the +system file or using an absolute name. + +For example: + +```text +mcpu-cpp -isystem $HOME/mcpu-wrapper ... +``` + +with `$HOME/mcpu-wrapper/math.h`: + +```text +#ifndef SOME_SYSTEM_MACRO +#define SOME_SYSTEM_MACRO temporary_value +#define REMOVE_SOME_SYSTEM_MACRO 1 +#endif + +#include_next <math.h> + +#ifdef REMOVE_SOME_SYSTEM_MACRO +#undef SOME_SYSTEM_MACRO +#undef REMOVE_SOME_SYSTEM_MACRO +#endif +``` + +If the home configuration also specifies: + +```text +MCPU_CPP_SYSTEM_INCLUDE_PATH = $HOME/mcpu-next/include; +``` + +a wrapper found through `-isystem` continues `#include_next` through the +configured user paths, then through `$HOME/mcpu-next/include/<lang>` and +`$HOME/mcpu-next/include`. The old system tree at the original installation +location does not participate. This is the intended way for a system developer +or tester to work in a private sandbox. + +MCPU-CPP stores the exact **physical element of the effective search chain** +from which the current header was found. `#include_next` starts at the next +element. The `"file"` and `<file>` forms of `#include_next` are equivalent; the +directory of the current file is not checked again. If the current file was +found by ordinary quoted search relative to its containing file and therefore +has no search-chain provenance, `#include_next` starts at the first element of +the configured chain. + +The operand may be produced by macro expansion. A logical name installed by +`#line` does not affect physical provenance. If no suitable file exists after +the current entry, preprocessing fails. + +### 4.5. `#pragma once` + +An active + +```text +#pragma once +``` + +directive marks the **physical file** as already processed during the current +MCPU-CPP invocation. A later attempt to include the same physical file skips +its contents. The directive itself is consumed by the preprocessor and is not +copied to output, including in `-dD` mode. + +Identity is determined by the file-system `st_dev`/`st_ino` pair, not by the +path string. Therefore the same file cannot bypass `#pragma once` by being +reached as `./file.h`, through a symbolic link, or through another hard-link +name. A logical name installed by `#line` also has no effect on this physical +identity. + +The mark takes effect immediately when the active directive is processed. +Therefore a header may include itself after `#pragma once`: the repeated +include is skipped and recursion does not occur. A directive in an inactive +conditional branch has no effect. + +MCPU-CPP recognizes only the exact `#pragma once` form, with optional +whitespace. Other `#pragma` directives are not interpreted by the preprocessor +and are preserved for later compiler stages; for example, `#pragma pack(...)` +continues to be passed through to output. + +`#pragma once` supplements, but does not modify, the normative +`#include`/`#include_next` search chain. The ordinary search mechanism first +finds a physical file, then the `once` registry decides whether its contents +must be processed. + +### 4.6. Forced files: `-imacros FILE` and `-include FILE` + +The command-line options + +```text +-imacros FILE +-include FILE +``` + +process a file before the primary input. They use the ordinary preprocessing +engine, not a separate simplified parser. + +The normative start-of-translation-unit order is: + +```text +predefined macros + -> -D/-U in command-line order + -> all -imacros in command-line order + -> all -include in command-line order + -> primary input +``` + +Thus the relative interleaving of `-imacros` and `-include` in `argv` does not +interleave the two groups: **all** `-imacros` files are always processed before +**all** `-include` files. + +`-imacros FILE` fully preprocesses the file. Its `#define`/`#undef`, +conditional directives, `#lang`/`#endlang`, `#include`, `#include_next`, +`#pragma once`, and diagnostics have normal semantics. However, all normal +preprocessing output from this forced file, including line markers and text +from nested headers, is discarded. The resulting macro-table state and other +preprocessing state are retained for later forced files and for the primary +input. + +`-include FILE` uses the same machinery, but preserves normal output, as if the +located header had been included immediately before the primary source. A +forced include is a real include boundary: inside it `__INCLUDE_LEVEL__ == 1`, +inside a header that it includes the level is `2`, and the primary input +remains at level `0`. `__BASE_FILE__` inside forced files remains the name of +the primary input. + +An absolute forced-file operand is used directly. A relative operand is first +searched for in the **current working directory**, then along the ordinary +include chain: + +```text +explicit -I +explicit -isystem +MCPU_CPP_<LANG>_INCLUDE_PATH +MCPU_CPP_INCLUDE_PATH +MCPU_CPP_SYSTEM_INCLUDE_PATH/<lang> +MCPU_CPP_SYSTEM_INCLUDE_PATH +explicit -idirafter +MCPU_CPP_AFTER_INCLUDE_PATH +``` + +The directory of the primary input receives no special priority when resolving +an `-imacros`/`-include` operand. Once a forced file is found, ordinary quoted +`#include "file"` inside it is again resolved relative to the physical +directory of that forced file. If the forced file was found through an element +of the include chain, its provenance is retained and `#include_next` continues +at the next chain element. + +Forced files and the headers actually reached from them participate in the +ordinary physical dependency registry. Their user/system classification is +derived from the same search provenance, so `-MM`/`-MMD` filter system forced +headers exactly as they filter ordinary system headers. A missing forced file +is an error. + +### 4.7. Dependency generation: `-M`, `-MM`, `-MG`, `-MD`, `-MMD`, `-MF`, `-MT`, `-MQ` + +`-M` and `-MM` use **the same include-pipeline pass** as ordinary preprocessing. +There is no second, independent header search. The dependency graph therefore +inherits the normative search order, `#include_next`, macro-expanded include +operands, conditional compilation, and `#pragma once` automatically. + +`-M` suppresses normal preprocessing output and emits one Make rule: + +```make +file.o: file.c header1.h header2.h +``` + +The list contains the primary source and all physical headers actually reached, +including system headers. A physical file appears only once. Identity is +`st_dev + st_ino`, so an alternate relative spelling, symbolic link, or hard +link does not create a duplicate dependency. A logical name established by +`#line` is only a source name and never enters the dependency list. + +`-MM` builds the same graph but removes system dependencies. System context +includes headers found through explicit `-isystem`, the configured system tree +`MCPU_CPP_SYSTEM_INCLUDE_PATH/<lang>` / `MCPU_CPP_SYSTEM_INCLUDE_PATH`, explicit +`-idirafter`, and `MCPU_CPP_AFTER_INCLUDE_PATH`, as well as the entire branch of +headers included directly or indirectly from such a system header. The syntax +`#include "file"` versus `#include <file>` does not by itself determine whether +a dependency is a system dependency. If one physical file is reached from a +system branch and is later included directly from user context, it remains a +user dependency and is present in `-MM` output. + +The default target is derived from the basename of the primary source: its +suffix is replaced with the object suffix (`.o` by default). Paths and the +default target are Make-quoted. For stdin, the GNU-like form is `-: -`. + +`-MD` and `-MMD` use the same dependency graph but, unlike `-M` and `-MM`, +**do not suppress normal preprocessing output**. `-MD` includes system headers +like `-M`; `-MMD` applies the user-only filter of `-MM`. One pass can therefore +produce both preprocessed text and a side-effect dependency file. + +If `-MF` is not specified, side-effect mode chooses the `.d` file name +automatically: + +* without `-o`, the input basename loses its suffix and receives `.d`; input + pathname directories are not copied into the dependency-file name; +* with an ordinary `-o FILE`, the output-file suffix is replaced by `.d`; +* stdin uses `-.d`. + +`-MF FILE` overrides the automatic dependency-file name. `-MF -` means stdout. +`-MF` also works with dependency-only `-M`/`-MM`; in that case it takes +priority over the ordinary destination of the Make rule. `-MF` by itself, +without one of `-M`, `-MM`, `-MD`, or `-MMD`, is an error. + +The semantics intentionally follow GNU CPP: `-MD`/`-MMD` do not accept their +own argument; `-MF` is a separate dependency-output option. + +`-MT TARGET` replaces the automatic target with `TARGET` **exactly as supplied**. +No Make quoting is performed. Thus one `-MT` argument may contain multiple +targets separated by spaces: + +```text +-MT 'obj/a.o obj/a.pic.o' +``` + +and repeated `-MT` options also append targets to the same rule: + +```text +-MT obj/a.o -MT obj/a.pic.o +``` + +`-MQ TARGET` has the same target-selection semantics but quotes characters that +are special to Make. For example: + +```text +-MQ '$(OBJDIR)/foo.o' +``` + +produces the left-hand side: + +```make +$$(OBJDIR)/foo.o: +``` + +Both separate arguments (`-MT TARGET`, `-MQ TARGET`) and attached forms +(`-MTTARGET`, `-MQTARGET`) are supported. + +If at least one `-MT` or `-MQ` is present, the automatic default target is not +emitted. In particular, `--object-suffix` affects only the automatic target and +does not rewrite explicit targets. When no explicit target is present, the +default target is Make-quoted as with `-MQ`. + +Repeated and mixed `-MT`/`-MQ` options are allowed. As in GNU CPP, all `-MT` +targets are emitted first in their command-line order, followed by all `-MQ` +targets in their command-line order. All of them form the left-hand side of +**one** dependency rule. + +`-MT` and `-MQ` are meaningful only with one of `-M`, `-MM`, `-MD`, or `-MMD`. +Using either without dependency generation is a command-line error. + +`-MG` changes only the handling of **missing** include files in dependency-only +`-M` and `-MM` modes. Without `-MG`, an unresolved `#include` remains an error. +With `-M -MG` or `-MM -MG`, a missing header is treated as a future generated +file: preprocessing does not fail and the directive operand is added to the +dependency rule **exactly as obtained after macro expansion**, without +prepending a guessed include directory. For example: + +```text +#include "generated.h" +``` + +adds `generated.h` under `-M -MG`, even when that file does not yet exist. A +macro-expanded include behaves the same way: the dependency receives the +expanded name. `-MG` is valid only with `-M` or `-MM`; combinations with +`-MD`/`-MMD`, or use without dependency-only mode, are command-line errors. + +Unresolved dependencies are integrated into **the same ordered dependency +registry** as physical files, but occupy a separate identity domain. For a +found file, the registry still uses `st_dev/st_ino` and physical provenance. +For a missing file there is no such information, so an `-MG` entry performs no +`stat()` and is deduplicated by the exact include-operand text. This is +essential: an identically named file in the current working directory must not +turn an unresolved `<name>` into a false physical match when angle search did +not find that file. Different unresolved spellings, such as `generated.h` and +`./generated.h`, are distinct dependencies. + +For `-MM`, an unresolved dependency gets its user/system class from search +context: a missing `<file>` is system-class and a missing `"file"` is +user-class when the including unit is not itself a system header; any missing +include reached from a system header remains system-class. If the same +unresolved operand appears more than once, the classification of its first +occurrence is retained, matching GNU CPP. Physical dependencies keep the +existing rule that if the same inode is later reached from user context, it is +no longer system-only. + +`-MG` also applies to missing command-line forced files `-include FILE` and +`-imacros FILE`: their operand enters the unresolved registry as a user +dependency without a synthetic search prefix. When the file exists, +`-include`/`-imacros` continue to use the normal physical dependency registry +and search provenance. + + +## 5. Language switching + +The preprocessor starts in language state `0`. This is the unnamed primary +C-like language and it is not a valid argument of `#lang`. + +Supported languages are: + +| Name | Purpose | +|---|---| +| `diff` | differential equations | +| `dift` | difference equations | +| `alg` | algebraic equations | +| `as` | MCPU assembler (`mcpu-as`) | +| `avm` | analog-computer schemes | +| `ACS` | block diagrams of automatic-control systems | + +`#lang` must be followed by a string constant containing one nonempty word: + +```text +#lang "diff" +``` + +The name is checked only against the internal language list above and is +compared case-insensitively in ASCII. Thus `"diff"`, `"Diff"`, `"DIFF"`, and +`"dIfF"` select the same language. The spelling inside the quotes is preserved +in output. + +Whitespace inside the string constant is forbidden: `" diff"`, `"diff "`, and +`"di ff"` are errors. Escape sequences are not interpreted inside this +constant. The closing quote must occur on the same physical source line. Only +whitespace is permitted between the closing quote and the end of the line. + +Whitespace outside the string is normalized. For example: + +```text + # lang "DiFf" +``` + +becomes: + +```text +#lang "DiFf" +``` + +`#lang` pushes a new language onto the language stack; `#endlang` restores the +previous language. The stack is not reset by `#include`, so a language block +may begin and end in different files. `#lang` and `#endlang` remain in the +output stream for the later frontend dispatcher; `#lang` is emitted in +normalized form. + +## 6. Object-like macro definitions + +Starting with 0.0.4, object-like macros are supported: + +```text +#define BUFFER_SIZE 1024 +#define NAME value +#define EMPTY +``` + +The `#define` directive itself is not copied to normal output. In ordinary +text, a macro identifier is replaced with its replacement list. That +replacement is rescanned for macro names, so cascaded expansion works: + +```text +#define A B +#define B 10 +A +``` + +produces `10`. + +While a particular macro is being expanded, that macro is temporarily +disabled. Therefore self-referential and mutually recursive definitions do not +cause infinite recursion. + +Macro names are not expanded inside string or character constants. For the +`diff` language, an apostrophe keeps its language-specific meaning and does +not protect following text as a C character constant. + +A multi-line definition using backslash-newline is supported because splicing +occurs before `#define` is parsed. + +### 6.1. `#undef` + +```text +#undef NAME +``` + +removes an object-like macro definition. Undefining a nonexistent macro is not +an error. + +### 6.2. Computed `#include` + +An `#include` argument that does not begin directly with `"` or `<` is first +macro-expanded. Therefore both of these are valid: + +```text +#define HEADER <diff/model.h> +#include HEADER +``` + +and + +```text +#define HEADER "local.h" +#include HEADER +``` + +The expansion result must have the form `"file"` or `<file>`. + +## 7. Function-like macros + +The macro engine supports function-like macros: + +```text +#define identifier( argument-list ) replacement +``` + +The opening parenthesis in the definition must follow the macro name +**immediately**. Thus + +```text +#define F(X) X +``` + +defines a function-like macro, while + +```text +#define F (X) +``` + +defines an object-like macro with replacement `(X)`. + +At a use site, whitespace is permitted between a function-like macro name and +the opening parenthesis. If `(` does not follow, the identifier is not a call +of that macro and remains in output. + +For an ordinary function-like macro, the number of actual arguments must equal +the number of formal parameters. For a variadic macro, every fixed argument +must be present, while the variadic tail may contain any number of arguments, +including an empty tail. Nested parentheses are tracked while parsing actual +arguments; a comma inside them does not separate arguments. Square brackets do +not have this property; this is part of the adopted macro-expansion semantics. + +For example: + +```text +#define min(X, Y) ((X) < (Y) ? (X) : (Y)) +min(1, 2) +``` + +produces: + +```text +((1) < (2) ? (1) : (2)) +``` + +Before substitution, an ordinary actual argument itself undergoes macro +expansion. Cascaded and nested calls therefore work naturally: + +```text +#define A 7 +#define min(X, Y) ((X) < (Y) ? (X) : (Y)) +min(min(A, 3), 10) +``` + +A formal parameter may occur any number of times in the replacement list. An +expression with side effects in an actual argument may therefore be evaluated +multiple times by the later compiler; the preprocessor does not attempt to +repair such source code. + +Macros with no formal parameters are supported: + +```text +#define READY() 1 +``` + +They expand only when called as `READY()` (whitespace between the name and `(` +is permitted at a use site); the standalone identifier `READY` does not +expand. + +Formal parameter names must be distinct. An unterminated parameter list, +invalid punctuation, and too few or too many actual arguments are errors. + +### 7.1. Stringification `#` + +The stringification operator (`#`) is supported for parameters of +function-like macros: + +```text +#define STR(X) #X +STR(alpha + beta) +``` + +produces: + +```text +"alpha + beta" +``` + +Stringification uses the **raw actual argument before macro expansion**. Thus: + +```text +#define A 7 +#define STR(X) #X +#define XSTR(X) STR(X) + +STR(A) -> "A" +XSTR(A) -> "7" +``` + +Leading and trailing whitespace in the argument is removed. Internal +whitespace sequences are collapsed to one space except inside string/character +tokens of the active language. Double quotes and backslashes inside quoted +tokens are escaped so that the result remains one valid string constant. + +In a function-like replacement list, `#` must refer, directly or after +whitespace, to a formal parameter name. Inside a quoted token, `#` is not an +operator. An empty actual argument is allowed and stringifies as `""`. + +### 7.2. Token concatenation `##` + +Starting with 0.0.21, token concatenation (`##`) is supported with semantics +aligned with GNU CPP and the macro engine's `collect_expansion()` / +`macroexpand()` model. The operator combines two adjacent preprocessing tokens +into one token, after which the resulting replacement list is rescanned for +macro expansion. + +For example: + +```text +#define CAT(A, B) A ## B +CAT(foo, bar) +``` + +produces `foobar`. Concatenation may form an identifier, preprocessing number, +or multi-character punctuator. For example: + +```text +CAT(1.5, e3) -> 1.5e3 +CAT(+, =) -> += +``` + +When a formal parameter is directly adjacent to `##`, its actual argument is +substituted **without preliminary macro expansion**. This is the same raw +argument principle used by stringification. To expand first and concatenate +second, use the normal two-level GNU CPP pattern: + +```text +#define AFTERX(X) X_ ## X +#define XAFTERX(X) AFTERX(X) +#define TABLESIZE 1024 +#define BUFSIZE TABLESIZE + +AFTERX(BUFSIZE) -> X_BUFSIZE +XAFTERX(BUFSIZE) -> X_1024 +``` + +An empty actual argument adjacent to `##` behaves as a placemarker: it adds no +token, and concatenation on that side leaves the remaining operand unchanged. +If an actual argument contains multiple preprocessing tokens, only the edge +token directly adjacent to `##` is concatenated; the others are preserved and +participate in the subsequent rescan. + +`#` and `##` may be used in the same function-like macro, for example: + +```text +#define COMMAND(NAME) #NAME | NAME ## _command +``` + +Here `#NAME` uses the raw spelling of the argument for stringification, while +`NAME ## _command` uses the same raw argument for concatenation. + +Inside a quoted token, `##` is not an operator. Comments have already become +whitespace by the time macro expansion occurs, so comments cannot be created +by concatenating `/` and `*`. Whitespace may originally appear between `##` +and its operands; it does not participate in the concatenation. + +If the two operands do not form one valid preprocessing token, a diagnostic is +issued and the original tokens are retained; whether whitespace appears +between them after that diagnostic is not part of the contract. `##` at the +beginning or end of a replacement list is a macro-definition error. + +### 7.3. Variadic macros: `...` and `__VA_ARGS__` + +Starting with 0.0.46, variadic function-like macros are supported in the modern +C99-compatible form: + +```text +#define LOG(...) output(__VA_ARGS__) +#define LOGF(format, ...) output(format, __VA_ARGS__) +``` + +The `...` marker may be the only parameter or the final element after one or +more fixed parameters. The old GNU extension with a named variadic parameter, + +```text +#define LOG(args...) ... +``` + +is intentionally not supported in 0.0.46. `__VA_OPT__` was also not part of +that particular release. + +At invocation, every token after the last fixed parameter, including commas +that separate those tokens, forms one logical variable argument and is +substituted for `__VA_ARGS__`. In an ordinary position, that variable argument +undergoes macro expansion before substitution, just like an ordinary actual +argument: + +```text +#define A 7 +#define V(...) <__VA_ARGS__> +#define F(first, ...) first | __VA_ARGS__ + +V(A, 2, 3) -> <7, 2, 3> +F(1, A, 3) -> 1 | 7, 3 +``` + +The variadic tail may be empty. Both + +```text +F(1) +F(1,) +``` + +are valid and substitute an empty `__VA_ARGS__`. This does **not** imply that a +comma written explicitly in the replacement list is removed automatically. +For example, with + +```text +#define E(format, ...) output(format, __VA_ARGS__) +``` + +`E("ok")` leaves the comma before the empty tail. The historical GNU +`, ## __VA_ARGS__` comma-swallowing behavior is deliberately outside the 0.0.46 +contract and remains unsupported; modern code should use `__VA_OPT__(,)`. + +`__VA_ARGS__` participates in the existing `#` and `##` semantics as a real +macro parameter. Stringification uses the raw spelling of the whole variadic +tail: + +```text +#define STRV(...) #__VA_ARGS__ +STRV(A, b + c) -> "A, b + c" +``` + +When adjacent to `##`, the variadic argument is likewise substituted without +prescan; the ordinary placemarker, token-concatenation, and rescan rules then +apply. For example: + +```text +#define L(...) pre ## __VA_ARGS__ +#define R(...) __VA_ARGS__ ## post + +L(fix) -> prefix +R(fix) -> fixpost +``` + +If the variadic argument contains multiple preprocessing tokens, only the edge +token immediately adjacent to `##` is concatenated and the remaining tokens +are preserved, exactly as for an ordinary parameter. An empty variadic tail +next to `##` behaves as a placemarker. + +The name `__VA_ARGS__` is reserved for the variable argument and is not +accepted as an ordinary formal parameter name. `#__VA_ARGS__` is valid only in +a variadic macro. Dump modes preserve the variadic form of the definition, for +example: + +```text +#define F(first,...) first | __VA_ARGS__ +``` + +### 7.4. `__VA_OPT__` + +Starting with 0.0.47, variadic macros support the standard conditional fragment +`__VA_OPT__(pp-tokens)`. If the variable argument contains no preprocessing +tokens after normal macro substitution, the entire `__VA_OPT__(...)` expands +to an empty sequence. If the variable argument is nonempty, the parenthesized +contents participate in the replacement list: + +```text +#define DEBUG(format, ...) \ + fprintf(stderr, format __VA_OPT__(,) __VA_ARGS__) + +DEBUG("ready") -> fprintf(stderr, "ready") +DEBUG("x=%d", x) -> fprintf(stderr, "x=%d", x) +``` + +Emptiness is decided **after expansion of the variable argument**, not from its +raw spelling. Therefore a macro that itself expands to an empty sequence does +not activate `__VA_OPT__`: + +```text +#define EMPTY +#define HAS(...) [__VA_OPT__(yes)] + +HAS() -> [] +HAS(EMPTY) -> [] +HAS(token) -> [yes] +``` + +The contents of `__VA_OPT__` may contain balanced nested parentheses. The +closing `)` of the `__VA_OPT__` construct is found with nesting taken into +account. A nested `__VA_OPT__` inside another `__VA_OPT__` is deliberately +forbidden. + +`__VA_OPT__` is integrated with the existing rules for parameter substitution, +stringification, token concatenation, placemarkers, and rescan. For example: + +```text +#define X 123 +#define S(...) #__VA_OPT__(__VA_ARGS__) +#define L(...) pre ## __VA_OPT__(__VA_ARGS__) + +S() -> "" +S(X) -> "123" +L() -> pre +L(X) -> pre123 +``` + +With `#__VA_OPT__(...)`, parameter substitution inside the fragment happens +first, including prescan of ordinary parameters, but arbitrary macro names in +the fragment are not additionally rescanned before stringification. Thus: + +```text +#define X 123 +#define S(a, ...) #__VA_OPT__(a X) + +S(X, y) -> "123 X" +``` + +If a parameter inside `__VA_OPT__` participates directly in an internal `##`, +prescan is suppressed for that parameter in the usual way; the paste is +performed before later rescan. An outer `##` adjacent to `__VA_OPT__` receives +the edge token of the already prepared fragment. An empty `__VA_OPT__` result +next to `##` behaves as a placemarker. + +`__VA_OPT__` is valid only in the replacement list of a variadic function-like +macro and must immediately introduce a parenthesized fragment. `##` cannot be +the first or last preprocessing token inside that fragment. + +The historical GNU extension + +```text +, ## __VA_ARGS__ +``` + +is intentionally **not implemented** by `mcpu-cpp`. Use the modern +`__VA_OPT__(,)` form for a conditional comma. The old GNU named variadic +parameter form `args...` also remains unsupported. + +### 7.5. Whitespace normalization in replacement lists + +Starting with 0.0.48, `mcpu-cpp` does not carry alignment whitespace from a +multi-line macro definition into the expansion result. After `\\` + newline +has been removed, a whitespace sequence belonging to the replacement list +itself is canonicalized to one ASCII space. This is particularly important for +definitions whose backslashes are visually aligned in one column: + +```text +#define TRACE(x) \ + do \ + { \ + output(x); \ + done(); \ + } \ + while( 0 ) +``` + +Such a definition expands to the compact replacement: + +```text +do { output(x); done(); } while( 0 ) +``` + +rather than preserving dozens of spaces before each former physical-line +boundary. + +Normalization applies **only to whitespace belonging to the replacement +list**. `mcpu-cpp` is not a source formatter: whitespace in ordinary input text +is preserved. Whitespace inside an actual macro argument is likewise not +reformatted merely because the argument is substituted into a macro: + +```text +#define ID(x) x + +ID(a + b) -> a + b +``` + +String and character literal contents are preserved verbatim, so: + +```text +#define S "left right" +``` + +still contains five spaces inside the string. + +The presence of whitespace between preprocessing tokens is preserved as one +space. This prevents accidental retokenization such as turning `+ +` into +`++`, `- >` into `->`, or `< <` into `<<`. The `#` and `##` operators, +placemarkers, `__VA_ARGS__`, `__VA_OPT__`, and later rescan keep their existing +rules; the policy changes only the amount of ordinary replacement-list +whitespace. + +Dump modes (`-dM`, `-dD`) show the same canonical replacement-list form stored +in the internal macro table. + +### 7.6. Invisible-line compaction and line markers + +Starting with 0.0.49, `mcpu-cpp` uses the same model as GNU CPP for vertical +whitespace: **remove it, but do not forget it**. Source lines that produce no +output preprocessing token after preprocessing need not remain as physical +blank lines in the `.E` output, but their source position still contributes to +line markers and to `__LINE__`. + +Why a line is invisible does not matter. It may be a consumed directive, an +inactive `#if` branch, a single-line or multi-line comment, an ordinary blank +line, or any mixture of these. The emitter compares its current output source +position with the position of the next line that will actually be emitted. + +If the next position is fewer than eight lines away, the gap is represented by +ordinary newlines. If the distance is eight lines or greater, the long run of +blank lines is replaced by a corrective line marker: + +```text +# N "file" +``` + +and the next content line immediately belongs to source line `N`. The behavior +therefore matches the GNU CPP boundary: gaps 0 through 7 use newlines; a gap of +8 or more uses a line marker. + +Structural enter/return markers for included files keep their ordinary +meaning: + +```text +# 1 "header.h" 1 +# 4 "source.c" 2 +``` + +If an included file produces no output, `mcpu-cpp` does not invent a marker +reporting how far the preprocessor progressed internally through that header. +An enter marker may be followed immediately by its return marker. The real +position is corrected again only when some following content must be emitted. + +This optimization changes only the representation of the output stream. +Source coordinates, `__LINE__`, diagnostics, `#line`, include enter/return +semantics, and macro processing remain tied to the logical source stream, not +to the number of physical lines in the compacted `.E` file. + +## 8. Predefined macros + +Starting with 0.0.6, the historical predefined-macro mechanism was restored in +the preprocessor. It is treated as a separate ABI/environment layer for the +future unnamed C-like language. These definitions are not decorative: their +names and values must match either GNU CPP semantics or an explicitly +documented MCPU/LibMPU contract. + +### 8.1. Dynamic source macros + +The following predefined macros are evaluated at the point of use: + +| Macro | Expansion | +|---|---| +| `__FILE__` | string constant containing the name of the current input file | +| `__LINE__` | decimal number of the current source line | +| `__BASE_FILE__` | string constant containing the primary input file name of the translation unit | +| `__INCLUDE_LEVEL__` | `#include` nesting level; `0` in the primary file | +| `__DATE__` | preprocessor start date in the form `"Mmm dd yyyy"` | +| `__TIME__` | preprocessor start time in the form `"hh:mm:ss"` | + +`__DATE__` and `__TIME__` share one timestamp for the whole translation unit. +Their special expansion is emitted without another macro rescan. + +These names reside in the ordinary macro table, so `#undef` followed by +`#define` may deliberately replace a builtin. + +### 8.2. Preprocessor version + +Starting with 0.0.8, the standalone preprocessor does not define GCC's +`__VERSION__`. That name belongs to a compiler environment, which does not yet +exist for the future high-level language. The version of `mcpu-cpp` has its +own unambiguous name: + +```text +#define __MCPU_CPP_VERSION__ "1.0.2" +``` + +The value is obtained automatically from `PACKAGE_VERSION`. When a compiler +frontend/driver appears, its version contract will be defined separately and +will not be mixed with the version of the standalone preprocessor. + +### 8.3. ABI sources of truth + +`mcpu-cpp` is built only with GNU GCC. During `configure`, the project follows +the established LibMPU/LibMPUIO `acsite.m4` approach: GCC predefined macros +describe native type sizes, byte/word order, and machine-register width, while +the installed `<libmpu.h>` is the final source of truth for LibMPU +configuration. + +In particular, the following values are captured and checked: + +```text +MPU_REAL_IO_LIMIT +MPU_MATH_FN_LIMIT +MPU_BYTE_ORDER +MPU_WORD_ORDER +BITS_PER_MACHINE_REGISTER +BITS_PER_UNIT_T +sizeof(__mpu_size_t) +sizeof(__mpu_ptrdiff_t) +``` + +`configure` additionally verifies that the byte order and +`BITS_PER_MACHINE_REGISTER` recorded by LibMPU agree with the GCC target used +to build `mcpu-cpp`. `MPU_WORD_ORDER` is taken directly from the configured +LibMPU profile and describes word order in the MCPU data environment. + +`MPU_REAL_IO_LIMIT` and `MPU_MATH_FN_LIMIT` serve different purposes. For +example, a library may support Real I/O up to 65536 bits while providing +mathematical functions only up to 16384 bits. Therefore `MPU_MATH_FN_LIMIT` +is not used as the limit on existence of Real types. + +### 8.4. MCPU architecture and assembler prefixes + +The target architecture is identified by: + +```text +#define _ARCH_MCPU 1 +``` + +MCPU PTR64 is 64 bits wide, so `__SIZEOF_POINTER__`, +`__MCPU_POINTER_WIDTH__`, `__INTPTR_TYPE__`, `__UINTPTR_TYPE__`, and the +corresponding width/max macros are defined accordingly. + +Assembler-prefix macros follow GNU CPP meaning rather than the first letter of +a register-view name. `mcpu-as` syntax uses no extra sigil before a register, +label, or immediate value. The letters `r` and `c` belong to MCPU register +syntax; they are not a `REGISTER_PREFIX`. Therefore: + +```text +#define __REGISTER_PREFIX__ +#define __LOCAL_LABEL_PREFIX__ +#define __USER_LABEL_PREFIX__ +#define __IMMEDIATE_PREFIX__ +``` + +all four expand to an empty sequence. `.L...` remains a compiler naming +convention and is not an assembler-ABI local-label prefix: LOCAL/GLOBAL binding +is determined by symbol directives. + +### 8.5. Byte order and word order + +The basic numeric byte-order values are compatible with GNU CPP: + +```text +__ORDER_LITTLE_ENDIAN__ +__ORDER_BIG_ENDIAN__ +__ORDER_PDP_ENDIAN__ +``` + +The target environment publishes its own MCPU names: + +```text +#define __MCPU_BYTE_ORDER__ __ORDER_LITTLE_ENDIAN__ +#define __MCPU_WORD_ORDER__ __ORDER_LITTLE_ENDIAN__ +#define __BYTE_ORDER__ __MCPU_BYTE_ORDER__ +``` + +The actual values of `__MCPU_BYTE_ORDER__` and `__MCPU_WORD_ORDER__` come from +the configured LibMPU profile (`MPU_BYTE_ORDER` and `MPU_WORD_ORDER`). They +therefore follow the host data representation for which LibMPU was built. This +does not alter the separate architectural contract for MCPU instruction +bytecode encoding. + +The GNU/C-specific name `__FLOAT_WORD_ORDER__` is not defined because the +future MCPU language has no `float` type. + +LibMPU/MCPU environment parameters are published in the MCPU namespace: + +```text +__MCPU_MACHINE_REGISTER_WIDTH__ +__MCPU_REAL_IO_LIMIT__ +__MCPU_MATH_FN_LIMIT__ +__MCPU_INT_MAX_WIDTH__ +__MCPU_REAL_MAX_WIDTH__ +__MCPU_COMPLEX_MAX_WIDTH__ +``` + +`__MCPU_INT_MAX_WIDTH__` is `NB_I_MAX * 8`; the Real/Complex maximum width is +the configured `MPU_REAL_IO_LIMIT`. `__MCPU_MACHINE_REGISTER_WIDTH__` is the +`BITS_PER_MACHINE_REGISTER` value of the installed LibMPU. Real I/O and math +limits are deliberately kept separate: `MPU_REAL_IO_LIMIT` controls existence +of Real/Complex type families and text conversion, while `MPU_MATH_FN_LIMIT` +controls availability of mathematical functions at a given width. + +### 8.6. MCPU size/ssize, `ptrdiff`, and pointers + +The future language does not inherit variable-width C names such as `short`, +`int`, and `long`, and it does not use the C-style name `size_t` as part of its +own ABI. The unsigned LibMPU size type and signed byte-count/error type are +published symmetrically in the MCPU namespace. For a 64-bit configured profile, +for example: + +```text +#define __MCPU_SIZE_TYPE__ uint64 +#define __MCPU_SIZE_WIDTH__ 64 +#define __MCPU_SIZEOF_SIZE__ 8 +#define __MCPU_SIZE_MAX__ 0xffffffffffffffff + +#define __MCPU_SSIZE_TYPE__ int64 +#define __MCPU_SSIZE_WIDTH__ 64 +#define __MCPU_SIZEOF_SSIZE__ 8 +#define __MCPU_SSIZE_MAX__ 0x7fffffffffffffff +``` + +This is an MCPU-specific family, not an attempt to invent a nonexistent GNU CPP +`__SSIZE_*` contract. + +The MCPU pointer ABI is independent of the host: PTR64 is always 64 bits wide: + +```text +#define __INTPTR_TYPE__ int64 +#define __UINTPTR_TYPE__ uint64 +#define __INTPTR_WIDTH__ 64 +#define __UINTPTR_WIDTH__ 64 +#define __INTPTR_MAX__ 0x7fffffffffffffff +#define __UINTPTR_MAX__ 0xffffffffffffffff +#define __SIZEOF_POINTER__ 8 +#define __MCPU_POINTER_WIDTH__ 64 +``` + +The difference between two MCPU pointers is signed and also fixed independently +of the host: + +```text +#define __PTRDIFF_TYPE__ int64 +#define __PTRDIFF_WIDTH__ 64 +#define __SIZEOF_PTRDIFF__ 8 +#define __PTRDIFF_MAX__ 0x7fffffffffffffff +``` + +Computed MIN expressions such as `(-__PTRDIFF_MAX__ - 1)` are not added to the +predefined table. + +### 8.7. Character types + +The future language has no ordinary C `char`. Therefore `__CHAR_TYPE__` and +`__WCHAR_TYPE__` are not defined. Language types are named without C/C++ `_t` +suffixes: + +```text +#define __CHAR8_TYPE__ char8 +#define __CHAR16_TYPE__ char16 +#define __CHAR8_WIDTH__ 8 +#define __CHAR16_WIDTH__ 16 +#define __SIZEOF_CHAR8__ 1 +#define __SIZEOF_CHAR16__ 2 +``` + +These are types of the future language. The implementation of `mcpu-cpp` +itself continues to use LibMPUIO `__mpu_char16_t` and the strict UCS-2 text +model internally. + +### 8.8. LibMPU integer families + +Complete structural metadata for integer families is generated up to the +actual `NB_I_MAX * 8` of the installed LibMPU rather than stopping at a +hard-coded final type. For every power-of-two width starting at 8 bits, TYPE, +WIDTH, and SIZEOF are defined: + +```text +#define __INT1024_TYPE__ int1024 +#define __UINT1024_TYPE__ uint1024 +#define __INT1024_WIDTH__ 1024 +#define __UINT1024_WIDTH__ 1024 +#define __SIZEOF_INT1024__ 128 +#define __SIZEOF_UINT1024__ 128 +``` + +With the current LibMPU 1.0.25, `NB_I_MAX == 8192`, so the family extends to +`int65536`/`uint65536`, with `__SIZEOF_INT65536__ == 8192`. + +Decimal-digit metadata is defined for **every** permitted integer width: + +```text +__INT<bits>_DECIMAL_DIG__ +__UINT<bits>_DECIMAL_DIG__ +``` + +The value is computed by `mcpu-cpp` integer-only helpers from the known bit +width. It is the exact number of decimal digits in the maximum value of the +type; neither a sign nor a terminating NUL is included in `DECIMAL_DIG`. For +unsigned values the maximum is `2^bits - 1`; for signed values it is +`2^(bits-1) - 1`. This differs from LibMPU `_int_digs()`, which estimates a +string-buffer size and includes room for a terminating NUL. + +For example: + +```text +#define __INT64_DECIMAL_DIG__ 19 +#define __UINT64_DECIMAL_DIG__ 20 +#define __INT256_DECIMAL_DIG__ 77 +#define __UINT256_DECIMAL_DIG__ 78 +``` + +Only the textual maxima themselves are deliberately limited to widths +`bits <= 256`: + +```text +__INT128_MAX__ +__UINT128_MAX__ +``` + +Maxima are produced through LibMPU `iuitoa()`. Macros named +`__INT<bits>_MIN__` are not generated: the predefined table must not contain +computed expressions such as `(-__INT<bits>_MAX__ - 1)`. For widths above 256 +bits, only MAX is absent; TYPE/WIDTH/SIZEOF/DECIMAL_DIG continue through the +full `NB_I_MAX * 8` range. + +### 8.9. LibMPU Real and Complex families + +Real/Complex structural metadata is generated for every power-of-two width +from 32 bits through the actual configured `MPU_REAL_IO_LIMIT`. TYPE, WIDTH, +and SIZEOF are published for all of these types. + +For Complex, WIDTH denotes the type parameter, not total storage width: + +```text +#define __COMPLEX128_TYPE__ complex128 +#define __COMPLEX128_WIDTH__ 128 +#define __SIZEOF_COMPLEX128__ 32 +``` + +`complex128` consists of two `real128` components, so its storage size is 32 +bytes. With `MPU_REAL_IO_LIMIT == 65536`, the top of the family is: + +```text +#define __COMPLEX65536_TYPE__ complex65536 +#define __COMPLEX65536_WIDTH__ 65536 +#define __SIZEOF_COMPLEX65536__ 16384 +``` + +For Real: + +```text +#define __REAL65536_TYPE__ real65536 +#define __REAL65536_WIDTH__ 65536 +#define __SIZEOF_REAL65536__ 8192 +``` + +Precision metadata is defined for **all** allowed Real widths up to +`MPU_REAL_IO_LIMIT`. Macro names correspond directly to LibMPU helpers: + +```text +__REAL<bits>_DECIMAL_DIG__ -> _real_digs(bits/8) +__REAL<bits>_MANT_DIG__ -> _real_mant_digs(bits/8) +``` + +`__REAL<bits>_DIG__` is intentionally absent. The `bits <= 256` restriction +applies only to large textual numeric constants. For widths up to 256 bits, +the following are also defined: + +```text +__REAL<bits>_MAX__ +__REAL<bits>_MIN__ +__REAL<bits>_EPSILON__ +__REAL<bits>_MAX_EXP__ +__REAL<bits>_MIN_EXP__ +__REAL<bits>_MAX_10_EXP__ +__REAL<bits>_MIN_10_EXP__ +``` + +For example, with LibMPU 1.0.25, the current `real128` profile gives values of +the form: + +```text +#define __REAL128_EPSILON__ 2.524354896707237777317531409e-29 +#define __REAL128_MAX__ 4.197157432934775384808581951e+323228496 +#define __REAL128_MIN__ 9.530259619551804292864984035e-323228497 +#define __REAL128_MAX_10_EXP__ 323228496 +#define __REAL128_MAX_EXP__ 1073741823 +#define __REAL128_MIN_10_EXP__ -323228524 +#define __REAL128_MIN_EXP__ -1073741822 +``` + +MAX/MIN/EPSILON are created by LibMPU itself and converted through +`real_to_ascii()`. Exponent constants are obtained from LibMPU exponent helpers +and integer conversion. For widths above 256 bits, these numeric predefines are +absent, but TYPE/WIDTH/SIZEOF/DECIMAL_DIG/MANT_DIG continue through +`MPU_REAL_IO_LIMIT`. + +For every supported Real type through `MPU_REAL_IO_LIMIT`, two compact +characteristics are also published: + +```text +#define __SIZEOF_REAL128_EXP__ 4 +#define __REAL128_MAX_STRLEN__ 60 +``` + +`__SIZEOF_REALxxx_EXP__` is obtained directly from `_sizeof_exp(NB_Rxxx)`. +`__REALxxx_MAX_STRLEN__` comes from `_real_max_string(NB_Rxxx)` and is the +maximum **number of characters** in the textual representation, not a byte +count. A zero-terminated string therefore needs at least +`__REALxxx_MAX_STRLEN__ + 1` elements: for `char8` that is the same number of +bytes, while for `char16` the physical byte count is twice as large. These two +metadata macros are also defined for Real widths above 256 bits because their +own values remain small. + +### 8.10. Macro dumps: `-dM`, `-dMP` + +The command: + +```text +mcpu-cpp -dM input.c +``` + +prints only final **non-predefined** macros in `#define ...` form. This group +includes definitions from the primary file and included headers, as well as +command-line `-D` definitions. MCPU-CPP's own predefined macros are not printed +by `-dM`. This mode is therefore intended primarily for a compact inspection of +macro state created by the user program. + +The command: + +```text +mcpu-cpp -dMP input.c +``` + +adds active MCPU-CPP predefined macros to the same final state. Output contains +two consecutive groups: predefined macros first, then non-predefined macros. +Definitions inside each group are sorted deterministically by name. This is +useful for system development because it exposes the preprocessing ABI and +architectural properties of the current MCPU environment without mixing them +with user definitions. + +Group membership is determined by macro origin, not by spelling. A macro +created by `-D` or `#define` is ordinary even if its name looks system-like. If +a predefined macro is removed with `#undef`, it is not printed. If the user +then defines the same name again, the new definition belongs to the ordinary +group and appears in the corresponding part of `-dMP`, and also in `-dM`. +Thus both modes display the **final macro state**. + +Context-dependent `__FILE__`, `__LINE__`, `__DATE__`, `__TIME__`, +`__BASE_FILE__`, and `__INCLUDE_LEVEL__` are not printed by the static dump. +Static ABI/architecture predefined macros and computed static Real metadata are +printed by `-dMP`. + +When an input file is supplied, it is fully preprocessed first and the final +macro state is printed afterward; ordinary preprocessed text is not emitted in +`-dM`/`-dMP` modes. Without an input file, stdin is used, so empty stdin with +`-dM` gives an empty dump while `-dMP` provides the active static predefined +macros of the current MCPU environment. + +`-dD` has different semantics and is unaffected by this distinction. + +### 8.11. Definition dump: `-dD` + +The command: + +```text +mcpu-cpp -dD input.c +``` + +preserves ordinary preprocessing output and additionally emits encountered +`#define` directives. Before primary input starts, static predefined macro +definitions are printed. Each such definition is preceded by a marker: + +```text +# 0 "<built-in>" +#define NAME value +``` + +and the predefined block itself is preceded by an input-file marker of the form +`# 0 "input.c"`. Context-dependent `__FILE__`, `__LINE__`, `__DATE__`, +`__TIME__`, `__BASE_FILE__`, and `__INCLUDE_LEVEL__` are not included in the +initial built-in block. + +### 8.12. Configuration dump: `-dconfig` + +The command: + +```text +mcpu-cpp -dconfig +``` + +requires no input file and prints the effective configuration-variable layer +after runtime configuration, optional system override, home user override, or +a selected `--config-file` have been read, including `$NAME`/`${NAME}` +expansion. Lines are sorted by name and printed as: + +```text +NAME = value; +``` + +This makes it possible to inspect actual include paths without manually +searching `<runtime-root>/etc/mcpu-cpp.conf`, `/etc/mcpu/mcpu-cpp.conf`, and +`$HOME/.mcpu/mcpu-cpp.conf`. + +### 8.13. Verbose configuration snapshot: `-v` + +With `-v`, MCPU-CPP retains its runtime trace for `#lang`, `#include`, and +`#include_next`, but configuration variables are printed only once, after all +configuration layers have been read and priority rules applied. Verbose output +therefore shows only **effective values**; intermediate values from runtime +root, system, and user configuration are not duplicated. + +The configuration block follows include-policy order: language-specific user +paths, the common user path, the system root, and the AFTER path. A variable +that is absent from every configuration layer is not printed. The +runtime-derived default `MCPU_CPP_SYSTEM_INCLUDE_PATH` is a full lowest-priority +value and is therefore visible under `-v` even when no `mcpu-cpp.conf` exists +**or all configuration files are disabled with `--no-config`**. + +The line form is: + +```text +config: NAME=value +``` + +### 8.14. Effective search directories: `-dsearch-dirs` + +The command: + +```text +mcpu-cpp -dsearch-dirs +``` + +requires no input file, prints the effective global search directories, and +exits without preprocessing. The format is intentionally simple: + +```text +search: /path/to/directory +``` + +Directories are printed in semantic search-class order: + +```text +explicit -I +explicit -isystem +configured language-specific user directories +MCPU_CPP_INCLUDE_PATH +MCPU_CPP_SYSTEM_INCLUDE_PATH/<lang> +MCPU_CPP_SYSTEM_INCLUDE_PATH +explicit -idirafter +MCPU_CPP_AFTER_INCLUDE_PATH +``` + +Language-specific entries are printed for every supported language in their +canonical order. During a real `#include`, only the directory corresponding to +the active `#lang` participates. The directory of the current physical file is +not printed by `-dsearch-dirs`: it exists only dynamically for a particular +`#include "..."` and changes with the include stack. `--no-config` does not +remove the runtime-derived system root, so even without configuration files the +dump still contains `<runtime-root>/include/<lang>` and +`<runtime-root>/include`. `-nostdinc` removes the effective system `<lang>` +entries and system root from the dump, but not explicit `-isystem`. A directory +that does not exist in the file system is still displayed because it remains +part of the effective search configuration and will simply be skipped during a +real file search. + +`-dsearch-dirs` accounts for `-I`, `-isystem`, `-idirafter`, every +configuration layer, and the replacement semantics of +`MCPU_CPP_SYSTEM_INCLUDE_PATH`. Using `-o` with this action is an error. + +### 8.15. Conditional compilation + +The directives `#if`, `#ifdef`, `#ifndef`, `#elif`, `#else`, and `#endif` are +processed as preprocessor control directives and are never copied to the +output stream, including under `-dD`. Inactive branches are skipped without +executing `#define`, `#undef`, or `#include` directives within them; nested +conditional groups are still tracked correctly. + +An `#if` expression first processes the `defined` operator, then undergoes macro +expansion, and any remaining identifiers evaluate to `0`. Arithmetic, +bitwise, comparison, and logical operators are supported, as are `?:` and +short-circuit semantics for `&&`, `||`, and `?:`. + +Starting with 0.0.26, expression syntax is parsed by a parser generated by ZUBR +4.1.0 from `src/mcpp-expr.zubr`; the same file contains the UCS-2 lexical +analyzer. `defined` preprocessing and macro expansion take place before the +parser is entered. Arithmetic semantics live in `mcpp-semantic.c/h` and do not +depend on the integer sizes of the host system. Generated `mcpp-expr.c` is +included in releases, so ZUBR is required only when the grammar changes. + +#### 8.15.1. The only evaluation width is 64 bits + +MCPU-CPP is a preprocessor, not a general-purpose language compiler. All +integer computation in conditional directives uses only 64-bit arithmetic. +The preprocessor does not perform arbitrary-width LibMPU arithmetic, floating +point, or complex-number computation. + +When a programmer does not need explicit control over the binary representation +of a literal, ordinary integer constants with optional `U`/`u` are sufficient. +For example: + +```c +#if 2 > 1 +#if 0xffffffffffffffffU > 1 +``` + +A numeric lexeme remains in UCS-2 until classification, after which its ASCII +portion is passed to LibMPU `iatoui()`. Binary `0b...`, octal `0...`, decimal, +and hexadecimal `0x...` forms are supported. A value that does not fit in 64 +bits is an error. Old C suffixes `L`, `l`, `LL`, and `ll` are not supported. + +#### 8.15.2. Width suffix `zNNN[Uu]` + +MCPU-CPP understands the width suffix shared by MCPU languages: + +```text +zNNN +ZNNN +zNNNu +zNNNU +ZNNNu +ZNNNU +``` + +`NNN` is a nonempty sequence of decimal digits and is **always** interpreted in +decimal, even with leading zeroes. Thus `z8`, `z08`, and `z008` all denote the +same width of 8 bits. + +In the general MCPU syntax, a valid width must be a power of two from 8 through +`MPU_REAL_IO_LIMIT`. MCPU-CPP, however, deliberately limits evaluation to 64 +bits: + +* `z8`, `z16`, `z32`, `z64`, in either letter case, are valid; +* `NNN > 64` is immediately an error: conditional preprocessing does not accept + numeric constants wider than 64 bits; +* if `NNN <= 64` but is not a valid power-of-two width, such as `z24`, a warning + is issued and the `zNNN` part itself is ignored; +* a following optional `U`/`u` selects unsigned interpretation and retains that + meaning even when an invalid `zNNN` has been ignored. + +The numeric preprocessing token must end after the complete suffix. An +operator or punctuation character begins the next token, so `1z32u+2`, +`(1z32u)`, and `1z32u==1` are valid. Forms such as `1z32undefined`, +`1z32ufoo`, and `1z32$foo` are errors and are not artificially split into a +number followed by a name. + +#### 8.15.3. Literal normalization + +The width suffix acts **exactly once, while the value of the literal itself is +formed**. The width is not retained in the semantic value and has no role in +later operations. + +For `VALUEzNNN`, the value is treated as a signed N-bit two's-complement number: + +1. retain the low `NNN` bits; +2. sign-extend the result to 64 bits. + +For `VALUEzNNNu`/`VALUEzNNNU`, the low `NNN` bits are retained and then +zero-extended to 64 bits. + +For example: + +```text +0x7fz8 -> 0x000000000000007f -> 127 +0x80z8 -> 0xffffffffffffff80 -> -128 +0xffz8 -> 0xffffffffffffffff -> -1 +0x80z8u -> 0x0000000000000080 -> 128 +0xffz8u -> 0x00000000000000ff -> 255 +0x1ffz8 -> 0xffffffffffffffff -> -1 +0x1ffz8u -> 0x00000000000000ff -> 255 +``` + +The last two examples are deliberate: `zNNN` specifies the width of the +**binary representation**, not a mathematical range check. Bits above N are +discarded before extension. + +After this normalization there is no remaining `z8`, `z16`, or `z32` concept +in the evaluation model. The internal value contains only a 64-bit bit pattern +and signed/unsigned state. + +#### 8.15.4. All subsequent operations are 64-bit + +After literal normalization, every arithmetic, bitwise, comparison, and logical +operation uses 64-bit operands. An operation result is not truncated back to +the width of the original suffix. Therefore: + +```text +0x7fz8 + 1 -> 128 +0xffz8u + 1 -> 256 +``` + +not `-128` and `0`. Likewise `~0xffz8u` inverts all 64 bits and gives +`0xffffffffffffff00`. + +For binary operations where signedness matters, the presence of an unsigned +operand selects 64-bit unsigned interpretation. Comparisons return `0` or `1`. +Logical `!`, `&&`, and `||` also return signed 64-bit `0` or `1`; short-circuit +evaluation does not evaluate an unselected operand. + +Shifts happen after 64-bit normalization. Right shift of a negative signed +value is arithmetic; right shift of an unsigned value is logical. For example: + +```text +0x80z8 >> 1 -> -64 +0x80z8u >> 1 -> 64 +``` + +The historical MCPU-CPP rule for a negative shift count is preserved: +`A << -N` is equivalent to `A >> N`, and `A >> -N` is equivalent to `A << N`. + +Thus `zNNN` does not turn the preprocessor into a compiler with integer +promotions over multiple widths. It only allows the binary representation of +the source literal to be stated explicitly; the expression then evaluates in +one simple 64-bit model. + +#### 8.15.5. Character constants + +A character unit has type `__mpu_uint16_t`, matching the internal UCS-2 +representation, and is zero-extended to 64 bits before evaluation. Subsequent +arithmetic is again ordinary 64-bit arithmetic. + +Conditional-compilation state is stored on a separate stack; a conditional +group may not cross an include-file boundary. + +### 8.16. Diagnostic directives `#error` and `#warning` + +MCPU-CPP supports the standard diagnostic directives: + +```text +#error message +#warning message +``` + +`#error` emits an error diagnostic using the current logical file name and line +number and immediately terminates preprocessing unsuccessfully. `#warning` +emits a warning with the same source-location information and preprocessing +continues. A preceding `#line` therefore affects both diagnostics. + +The remainder of the line after the directive name **does not undergo macro +expansion**. For example: + +```c +#define MESSAGE expanded +#warning MESSAGE +``` + +prints `MESSAGE`, not `expanded`. This distinguishes diagnostic directives from +`#if` and `#line`, where macro expansion is part of the relevant contract. + +Comments are removed by the ordinary preprocessing phase before the directive +is processed. Leading and trailing whitespace in the message is removed and +whitespace sequences between preprocessing tokens are collapsed to one space. +Whitespace inside quotes is preserved. For example: + +```c +#warning one /* comment */ two +#warning "a b" +``` + +produce `one two` and `"a b"`, respectively. Unicode text passes through the +internal UCS-2 representation and is written to the external diagnostic as +UTF-8. + +Both directives are control directives and are never copied to normal output +or to `-dD`. In an inactive `#if` branch they are ignored completely, so the +usual protective pattern behaves as expected: + +```c +#if 0 +#error this error is inactive +#endif +``` + +### 8.17. Warning control: `-Wcomment`, `-Wall`, `-Werror` + +MCPU-CPP distinguishes mandatory warnings that are part of established +preprocessing semantics from optional warning classes enabled by the user. +Warning control does not alter `-dD`, macro expansion, conditional compilation, +or include search semantics. + +`-Wcomment` and `-Wcomments` are exact aliases and enable two lexical warnings: + +* a `/*` sequence seen while already inside an open `/* ... */` comment; +* backslash-newline inside a `//` comment, causing that single-line comment to + continue physically onto the next source line. + +This optional class is disabled by default. `-Wall` enables all optional +MCPU-CPP warning classes; in version 0.0.40 this class is `-Wcomment`. +`-Wno-comment` and `-Wno-comments` disable it. As in the GNU warning model, a +more specific setting has priority over a group setting regardless of argument +order. Therefore both: + +```text +mcpu-cpp -Wall -Wno-comment file.c +mcpu-cpp -Wno-comment -Wall file.c +``` + +leave comment warnings disabled. Between settings of equal specificity, the +last option wins; for example `-Wno-comment -Wcomment` enables the class. + +`-Werror` does not enable any new warning class. It promotes to an error every +warning that would actually be emitted during that invocation, causing an +unsuccessful result. This applies both to optional comment warnings and to +existing mandatory MCPU-CPP warnings, including: + +* an active `#warning` directive; +* an invalid `zNNN` width not exceeding 64 bits; +* redefinition of a macro with a different replacement list; +* a `##` result that does not form a single preprocessing token. + +For example: + +```text +mcpu-cpp -Wcomment -Werror file.c +``` + +turns a detected comment warning into an error. By contrast, `-Werror` alone, +without `-Wcomment`/`-Wall`, does not cause MCPU-CPP to search for optional +comment warnings. + +`-Wno-error` restores ordinary warning severity. Between `-Werror` and +`-Wno-error`, which have the same specificity, the last command-line option +wins. Thus `-Werror -Wno-error` leaves warnings as warnings, while +`-Wno-error -Werror` promotes them again. + +Version 0.0.40 deliberately did not introduce `-Werror=<class>`, +`-Wno-error=<class>`, `-Wundef`, `-Wunused-macros`, `-Wtraditional`, or other +compiler-oriented classes. The MCPU-CPP warning interface remains compact and +is extended only when a class is actually needed by the preprocessing +language itself. + +### 8.18. UCS-2 identifiers + +Starting with 0.0.22, preprocessing identifiers are no longer restricted to +ASCII. Inside `mcpu-cpp`, text is already strict UCS-2, and characters are +classified by locale-independent LibMPUIO 1.0.4 functions based on Unicode +18.0.0. The first identifier character must be `_` or have the `XID_Start` +property; following characters must be `_`, `$`, or have `XID_Continue`. +`$` is an `mcpu-cpp` extension: it is allowed only after the first character +and may not start an identifier. One rule is used consistently for macro names +and parameters, `#undef`, `#ifdef`/`#ifndef`, `defined`, ordinary macro +expansion, and `#`/`##`. Names remain case-sensitive. Surrogate code units +`U+D800..U+DFFF` are not valid identifier characters. + +For example, all of these are valid: + +```c +#define АНДРЕЙ 1 +#define résumé 2 +#define ΩМЕГА 3 +#define VALUE$OLD 4 +``` + +`VALUE$OLD` is valid, while `$VALUE` is invalid because `$` is not an +identifier-start character. + +Combining marks and non-ASCII decimal digits may appear in `XID_Continue` +positions but do not automatically become valid initial characters. Numeric +constant syntax is unaffected: it follows the rules of the active language, +not Unicode `isdigit`. + +### 8.19. Command-line macros `-D` and `-U` + +Starting with 0.0.23, `-D` and `-U` are full preprocessor actions. Supported +forms are: + +```text +-DNAME +-DNAME=VALUE +-D'FUNC(a,b)=a+b' +-UNAME +``` + +`-DNAME` is equivalent to `#define NAME 1`; an `=` with an empty right-hand +side defines an empty replacement list. Function-like command-line definitions +use the same macro engine as ordinary `#define`, including parameters, `#`, +`##`, and subsequent rescanning. `-U` uses the same identifier contract as +`#undef`. `-D`/`-U` actions are executed in command-line order after predefined +macros have been installed. + +Only the payload of `-D` and `-U` is interpreted as UTF-8 and converted to +strict UCS-2. File names, `-I`, other pathname arguments, and all other +command-line arguments remain the original byte strings and undergo no Unicode +conversion. + +Starting with 0.0.25, `$` is allowed inside a macro name, but not in its first +position. When `$` is passed through a shell, the user must account for shell +rules: the shell processes `$` **before `mcpu-cpp` starts**. Single quotes +fully protect `$`, for example: + +```sh +mcpu-cpp '-DАНДРЕЙ$_Y=62' input.c +``` + +Without quotes, `$` must be escaped: + +```sh +mcpu-cpp -DАНДРЕЙ\$_Y=62 input.c +``` + +or double quotes may be used with escaping: + +```sh +mcpu-cpp -D"АНДРЕЙ\$_Y=62" input.c +``` + +The unprotected form: + +```sh +mcpu-cpp -DАНДРЕЙ$_Y=62 input.c +``` + +does not pass the spelling literally: `$...` is expanded by the shell first, +and `mcpu-cpp` receives the already modified `argv`. Inside single quotes, a +backslash before `$` is unnecessary and would become an ordinary argument +character. + +For command-line `-D`, the left-hand side up to the first `=` is parsed as a +separate macro declarator. If an invalid tail occurs after a valid name (or a +completed formal-parameter list of a function-like macro) but before `=`, that +tail is silently discarded and **never becomes part of the replacement list**. +For example: + +```text +-D'АНДРЕЙ@XYZ=62' +``` + +is equivalent to: + +```c +#define АНДРЕЙ 62 +``` + +not the invalid `#define АНДРЕЙ @XYZ 62`. The same valid-identifier-prefix rule +applies to `-U`. If the very first character is not a valid identifier-start +character, such as `$` or a digit, the definition remains an error. + +Under `-dD`, definitions originating from `-D` are marked separately from +predefined macros: + +```text +# 0 "<command-line>" +#define NAME value +``` + +while predefined macros continue to use `<built-in>`. + +### 8.20. Public command-line interface + +`mcpu-cpp` supports only the current options documented by `--help`. Obsolete +compatibility flags do not form a hidden interface and are diagnosed as +`unknown option`. `-E` is the exception: it is silently accepted and ignored +because a compiler driver may pass it while invoking a standalone +preprocessor. + +`--object-suffix SUFFIX` selects the object-target suffix used when generating +Make dependencies; the argument is mandatory. + +## 9. Build system and generators + +Project-owned Autoconf macros live in the root `acsite.m4`. The `m4/` directory +is reserved for external/vendor M4 files. This follows the convention used by +MCPU libraries and keeps project configure code separate from imported macros. + +The `#if` expression parser is generated by ZUBR 4.1.0 from +`src/mcpp-expr.zubr`. Release archives contain both the grammar and the already +generated `src/mcpp-expr.c`, so an ordinary release build does not require +ZUBR. After the grammar changes, a developer build uses the normal Automake +rule: + +```text +zubr -vl -s -Bmcpp_ -o mcpp-expr.c mcpp-expr.zubr +``` + +Before a release, generated C must correspond to the grammar, and the full test +suite plus `make distcheck` must complete without errors. + +### 9.1. Developer bootstrap and the Git source tree + +Starting with 0.0.50, the root `./bootstrap` script makes it unnecessary to +store files in Git when they are completely reproducible from source. The +script first generates `src/mcpp-expr.c` from `src/mcpp-expr.zubr` using ZUBR +4.1.0, then runs `aclocal`, `autoheader`, `automake`, and `autoconf` in the +style used by LibMPU and LibMPUIO. `--target-dest-dir=DIR` selects a target +ROOTFS used for the system Autoconf macro/include directories. + +This convention applies specifically to the developer Git tree. **Release +archives remain self-contained**, exactly as before: they contain `configure`, +`Makefile.in`, Automake helper scripts, and the generated `src/mcpp-expr.c`. +Therefore an ordinary release build requires neither a preliminary `bootstrap` +run nor ZUBR. + +The root `.gitignore` lists reproducible bootstrap files and ordinary +configure/build state. It does not change the existing release/build model; it +only allows a cleaner Git repository. + +## 10. GNU-compatible features + +`mcpu-cpp` is an independent MCPU preprocessor, but it intentionally follows +GNU CPP behavior for a number of well-known operations. Compatibility applies +to the documented features; it does not imply complete CLI or language +interchangeability with GCC. + +GNU-compatible behavior is used for, in particular: + +* object-like and function-like macros, macro rescan, `#`, and `##`; +* variadic macros `...` / `__VA_ARGS__` and standard `__VA_OPT__`; +* `#if`, `#ifdef`, `#ifndef`, `#elif`, `#else`, `#endif`, and `defined`; +* `#include`, `#include_next`, `#pragma once`, `#line`, and GNU linemarkers; +* compact output mapping: up to seven invisible lines are represented by + newlines, while a gap of eight or more uses a corrective linemarker; +* forced files `-include` / `-imacros` and dependency options `-M`, `-MM`, + `-MD`, `-MMD`, `-MF`, `-MT`, `-MQ`, and `-MG`; +* warning controls `-w`, `-Wall`, `-Werror`, and the supported `-Wcomment` forms. + +MCPU-specific facilities, including `#lang` / `#endlang`, the `zNNN` numeric +suffix, and ABI predefined macros, remain native `mcpu-cpp` extensions. diff --git a/doc/mcpu-cpp-ru.md b/doc/mcpu-cpp-ru.md new file mode 100644 index 0000000..f8f762a --- /dev/null +++ b/doc/mcpu-cpp-ru.md @@ -0,0 +1,2167 @@ +# mcpu-cpp + +`mcpu-cpp` — препроцессор языков программирования MCPU. Он является +самостоятельным компонентом экосистемы LibMPU/LibMPUIO/LibMCPU и не привязан +к названию одного конкретного языка: активный язык выбирается директивой +`#lang`. + +Этот документ задаёт нормативное поведение `mcpu-cpp`: текстовую модель, +директивы, macro engine, include pipeline, конфигурацию, диагностику и +генерацию зависимостей для инструментов MCPU. + +## 1. Текстовая модель + +Внешние исходные файлы и конфигурационные файлы имеют кодировку UTF-8. +UTF-8 должен быть корректным. Для исходных программ проверка принадлежности +символов диапазону UCS-2 выполняется после удаления комментариев: поэтому +корректный Unicode scalar value выше `U+FFFF` допустим внутри комментария, но +остаётся ошибкой в программном тексте. После этой стадии исходный текст +обрабатывается как последовательность `__mpu_char16_t`. Входной UTF-8 BOM +допускается и удаляется. Встроенный NUL в исходном файле запрещён. + +Переводы строк `CRLF` и `CR` нормализуются в `LF`. + +## 2. Действия, выполняемые независимо от директив + +`mcpu-cpp` выполняет несколько преобразований до разбора +директив. + +### 2.1. Backslash-newline + +Последовательность `\\` непосредственно перед переводом строки удаляется до +распознавания комментариев, директив и макросов. Поэтому, например, + +```text +#defi\ +ne FOO 10\ +20 +``` + +эквивалентно логической строке + +```text +#define FOO 1020 +``` + +При этом физические номера строк продолжают учитываться при формировании +текущей позиции; если пользователь не менял её директивой `#line`, они и будут +видны в генерируемых line marker-ах. + +### 2.2. Комментарии + +Комментарии `/* ... */` и `// ...` удаляются до последующей обработки. Там, +где это необходимо для разделения соседних токенов, сохраняется пробельный +разделитель. Если комментарий завершает непустую строку, после его удаления не +сохраняются ни синтетический разделитель, ни пробелы, предшествовавшие +комментарию: строка заканчивается последним значащим символом. Это относится и +к многострочному комментарию, начавшемуся после программного текста. Если после +удаления комментария строка вообще не содержит ничего кроме пробелов, она +становится действительно пустой строкой. При этом комментарий между двумя +токенами по-прежнему оставляет необходимый разделитель и не склеивает их. +Переводы строк сохраняются, чтобы не разрушать координаты исходного текста. + +Комментарий не распознаётся внутри строковой или символьной константы. Для +языка `diff` апостроф не считается началом символьной константы, поскольку +используется в обозначениях производных. + +В буквальном аргументе `#include <...>` последовательности `/*` и `//` +рассматриваются как часть имени файла. + +## 3. Директивы и выходной поток + +Директива начинается символом `#`, если до него в логической строке находятся +только пробельные символы или комментарии. Между `#` и именем директивы +допускаются пробелы. + +Служебная информация о позиции в выходном потоке представлена в форме GNU +**line marker**: + +```text +# номер "имя-файла" [флаги] +``` + +Это не входная директива `#line`. При входе во включаемый файл к line marker-у +добавляется флаг `1`, а при возврате в файл, содержащий `#include`, — флаг `2`. +Эти значения имеют тот же смысл, что и в GNU CPP: `1` означает вход в новый +файл, `2` — возврат в предыдущий файл. Флаг `2` не является числом или уровнем +вложенности. + +Например: + +```text +# 1 "main.c" +# 1 "defs.h" 1 +... +# 2 "main.c" 2 +``` + +Входная директива + +```text +#line 62 "main.y" +``` + +сама в выходной поток не копируется. Она изменяет логические значения +`__LINE__` и `__FILE__` для последующего текста, а в выходе представляется +line marker-ом: + +```text +# 62 "main.y" +``` + +Аргументы `#line` предварительно подвергаются macro expansion, как в +принятой модели line control. Если после такого `#line` происходит `#include`, то после +возврата marker получает флаг `2`, например `# 65 "main.y" 2`. Имя, заданное +через `#line`, становится логическим именем для `__FILE__` и line marker-ов; оно +не меняет каталог, относительно которого ищется quoted `#include`. + +Директивы препроцессора имеют только канонические английские имена. Unicode остаётся полностью допустимым в идентификаторах, строках, комментариях и другом пользовательском тексте. + +## 4. Заголовочные файлы + +Поддерживаются: + +```text +#include "file" +#include <file> +#include_next "file" +#include_next <file> +#pragma once +``` + +Для обычного `#include "file"` первым всегда проверяется каталог **физического** +текущего исходного файла. Логическое имя, установленное через `#line`, на этот +шаг не влияет. Для `#include <file>` каталог текущего файла не проверяется. + +### 4.1. Перемещаемый корень MCPU как общий принцип экосистемы + +Начиная с выпуска 0.0.37 каталог установки MCPU **не содержит версию +конкретного инструмента** и не является абсолютной runtime-константой, +зашитой в бинарный файл. Версия относится к самому `mcpu-cpp`, `mcpu-as`, +`mcpu-ld`, `mcpu-run` или библиотеке, но не определяет корень единой среды +MCPU. + +При типичной конфигурации: + +```text +./configure --prefix=/usr --libdir=/usr/lib64 +``` + +`make install` создаёт дерево: + +```text +/usr/lib64/mcpu/ +├── bin/ +│ └── mcpu-cpp +├── etc/ +│ └── mcpu-cpp.conf +├── include/ +│ ├── diff/ +│ ├── dift/ +│ ├── alg/ +│ ├── as/ +│ ├── avm/ +│ └── acs/ +└── lib/ # общий каталог будущих библиотек MCPU +``` + +Публичное имя программы находится в `$bindir`: + +```text +/usr/bin/mcpu-cpp -> ../lib64/mcpu/bin/mcpu-cpp +``` + +Абсолютный `/usr/lib64/mcpu` при этом **не является частью runtime ABI +MCPU-CPP**. Он используется только `make install` как выбранное configure-time +место размещения файлов. + +При каждом обычном запуске MCPU-CPP определяет фактический путь собственного +исполняемого файла через Linux `/proc/self/exe`. Символическая ссылка публичной +команды не мешает этому: `/proc/self/exe` указывает на реально выполняемый +бинарный файл. Если `/proc/self/exe` недоступен, используется резервное +разрешение `argv[0]` через `PATH` и `realpath(3)`; возврата к зашитому +configure-time installation root нет. + +Для бинарного файла: + +```text +<root>/bin/mcpu-cpp +``` + +runtime-корень определяется как: + +```text +executable = <root>/bin/mcpu-cpp +executable dir = <root>/bin +MCPU runtime root = <root> +``` + +Из него автоматически выводятся: + +```text +<root>/etc/mcpu-cpp.conf +<root>/include +``` + +Следовательно всё дерево можно физически перенести, например из: + +```text +/usr/lib64/mcpu/ +``` + +в: + +```text +/opt/mcpu-test/ +``` + +или: + +```text +$HOME/devel/mcpu-next/ +``` + +и `<new-root>/bin/mcpu-cpp` без переконфигурирования начнёт использовать +`<new-root>/etc/mcpu-cpp.conf` и `<new-root>/include`. Старый абсолютный путь +не сохраняется ни в runtime default, ни в штатном `mcpu-cpp.conf`. + +Это не частная особенность препроцессора, а **общий принцип экосистемы MCPU**. +Будущие `mcpu-as`, `mcpu-ld`, `mcpu-run`, библиотеки, CRT и другие компоненты +должны разделять один перемещаемый корень: + +```text +<root>/bin +<root>/etc +<root>/include +<root>/lib +``` + +Их собственные версии могут отличаться, но согласованность конкретной среды +MCPU определяется тем, что все компоненты находятся в одном runtime tree, а +не совпадением version suffix в именах каталогов. + +### 4.2. Runtime defaults, уровни конфигурации и системный include root + +До чтения любого конфигурационного файла MCPU-CPP создаёт runtime-derived +значение: + +```text +MCPU_CPP_SYSTEM_INCLUDE_PATH = <runtime-root>/include +``` + +После этого конфигурационные слои применяются в порядке возрастающего +приоритета: + +```text +runtime-derived defaults + ↓ +<runtime-root>/etc/mcpu-cpp.conf + ↓ +/etc/mcpu/mcpu-cpp.conf + ↓ +$HOME/.mcpu/mcpu-cpp.conf +``` + +`<runtime-root>/etc/mcpu-cpp.conf` устанавливается вместе с MCPU-CPP, но сам +файл намеренно не содержит абсолютного штатного `MCPU_CPP_SYSTEM_INCLUDE_PATH`: +иначе перенос всего дерева восстановил бы старый путь. `/etc/mcpu/mcpu-cpp.conf` +является необязательным machine-wide override: `make install` каталог +`/etc/mcpu` не создаёт. Домашний `$HOME/.mcpu/mcpu-cpp.conf` также необязателен, +не версионируется и имеет максимальный config-приоритет. + +Если одна переменная определена несколько раз, побеждает последнее +определение, включая пустое. Поэтому `MCPU_CPP_SYSTEM_INCLUDE_PATH` остаётся +полностью заменяемым system root. Например: + +```text +MCPU_CPP_SYSTEM_INCLUDE_PATH = $HOME/mcpu-next/include; +``` + +полностью заменяет runtime-derived `<runtime-root>/include`. Для активного +`#lang "as"` тогда проверяются: + +```text +$HOME/mcpu-next/include/as +$HOME/mcpu-next/include +``` + +Штатные language-подкаталоги всегда выводятся самим препроцессором из одного +root; переменных вида `MCPU_CPP_SYSTEM_<LANG>_INCLUDE_PATH` нет. + +Пустое effective значение: + +```text +MCPU_CPP_SYSTEM_INCLUDE_PATH = ; +``` + +удаляет configured system stage полностью. Более приоритетный config может +после этого снова включить его непустым значением. + +`--config-file FILE` применяет явно выбранный файл поверх runtime-derived +default. `--no-config` отключает **только чтение файлов конфигурации**: +`<runtime-root>/etc/mcpu-cpp.conf`, `/etc/mcpu/mcpu-cpp.conf` и +`$HOME/.mcpu/mcpu-cpp.conf` не читаются, но `<runtime-root>/include` остаётся +штатным system root. Только `-nostdinc` удаляет effective standard-system tree +из include search для конкретного запуска; явно переданный `-isystem` при этом +остаётся command-line каталогом. + +### 4.3. Нормативный порядок поиска include-файлов + +Порядок поиска является частью контракта MCPU-CPP. Явно заданные параметры +командной строки имеют приоритет над persistent configuration. После +необязательного каталога текущего физического файла эффективная цепочка имеет +строго следующий вид: + +```text +explicit -I + ↓ +explicit -isystem + ↓ +MCPU_CPP_<LANG>_INCLUDE_PATH + ↓ +MCPU_CPP_INCLUDE_PATH + ↓ +MCPU_CPP_SYSTEM_INCLUDE_PATH/<lang> + ↓ +MCPU_CPP_SYSTEM_INCLUDE_PATH + ↓ +explicit -idirafter + ↓ +MCPU_CPP_AFTER_INCLUDE_PATH +``` + +Элементы, которых нет или которые не содержат требуемого файла, пропускаются. + +`MCPU_CPP_<LANG>_INCLUDE_PATH` — свободно настраиваемые пользователем +language-specific path-list'ы: + +```text +MCPU_CPP_DIFF_INCLUDE_PATH +MCPU_CPP_DIFT_INCLUDE_PATH +MCPU_CPP_ALG_INCLUDE_PATH +MCPU_CPP_AS_INCLUDE_PATH +MCPU_CPP_AVM_INCLUDE_PATH +MCPU_CPP_ACS_INCLUDE_PATH +``` + +Пользователь полностью распоряжается именами и расположением этих каталогов. +`MCPU_CPP_INCLUDE_PATH` — общий пользовательский path-list, видимый во всех +языковых состояниях. + +`-idirafter` и `MCPU_CPP_AFTER_INCLUDE_PATH` являются общим fallback-карманом. +MCPU-CPP не строит для них автоматических `<lang>`-подкаталогов. Пользователь +сам организует их внутреннюю структуру и при необходимости пишет, например: + +```text +#include <vendor/device.h> +``` + +Именно semantic class, а не порядок появления разных классов в argv/config, +определяет приоритет. Внутри одного класса сохраняется порядок добавления. + +### 4.4. `#include_next` и wrapper headers + +`#include_next` предназначен прежде всего для заголовков-обёрток (wrapper +headers). Он позволяет поставить локальный header раньше системного, изменить +локальную политику и затем продолжить поиск одноимённого header по нормативной +цепочке без копирования системного файла и без абсолютного имени. + +Например: + +```text +mcpu-cpp -isystem $HOME/mcpu-wrapper ... +``` + +и `$HOME/mcpu-wrapper/math.h`: + +```text +#ifndef SOME_SYSTEM_MACRO +#define SOME_SYSTEM_MACRO temporary_value +#define REMOVE_SOME_SYSTEM_MACRO 1 +#endif + +#include_next <math.h> + +#ifdef REMOVE_SOME_SYSTEM_MACRO +#undef SOME_SYSTEM_MACRO +#undef REMOVE_SOME_SYSTEM_MACRO +#endif +``` + +Если домашний config одновременно задаёт: + +```text +MCPU_CPP_SYSTEM_INCLUDE_PATH = $HOME/mcpu-next/include; +``` + +wrapper найденный через `-isystem` продолжит `#include_next` уже через +configured user paths, затем через +`$HOME/mcpu-next/include/<lang>` и `$HOME/mcpu-next/include`; старое system tree исходного места установки при этом не участвует. Именно такой сценарий позволяет +системному разработчику или тестеру жить в собственной sandbox. + +MCPU-CPP хранит конкретный **физический элемент effective search chain**, из +которого найден текущий header. `#include_next` начинает со следующего элемента. +Формы `"file"` и `<file>` для `#include_next` эквивалентны; каталог текущего +файла повторно не проверяется. Если текущий файл найден обычным quoted-поиском +относительно содержащего файла и не имеет search-chain provenance, +`#include_next` начинает с первого элемента configured chain. + +Операнд может быть получен macro expansion. Логическое имя после `#line` не +влияет на физический provenance. Если после текущего entry подходящего файла +нет, preprocessing завершается ошибкой. + +### 4.5. `#pragma once` + +Активная директива + +```text +#pragma once +``` + +помечает **физический файл** как уже обработанный в текущем запуске +MCPU-CPP. При последующей попытке включить тот же физический файл его +содержимое повторно не обрабатывается. Сама директива потребляется +препроцессором и в выходной поток не копируется, в том числе при `-dD`. + +Идентичность определяется по паре `st_dev`/`st_ino`, полученной файловой +системой, а не по строковому имени пути. Поэтому один и тот же файл не может +обойти `#pragma once`, если он достигнут как `./file.h`, через символическую +ссылку или через другое жёсткое имя (hard link). Логическое имя после `#line` +также не влияет на эту физическую идентичность. + +Пометка действует сразу в момент обработки активной директивы. Поэтому +заголовок может после `#pragma once` включить самого себя: повторное включение +будет пропущено и рекурсия не возникнет. Директива внутри неактивной ветви +условной компиляции никакого действия не имеет. + +MCPU-CPP распознаёт только точную форму `#pragma once` с необязательными +пробелами. Остальные `#pragma` не интерпретируются препроцессором и сохраняются +для последующих стадий компиляции; например, `#pragma pack(...)` продолжает +передаваться в выходной поток. + +`#pragma once` дополняет, но не изменяет нормативную search-chain +`#include`/`#include_next`: сначала обычный механизм поиска находит физический +файл, затем registry `once` решает, надо ли обрабатывать его содержимое. + +### 4.6. Принудительные файлы: `-imacros FILE` и `-include FILE` + +Опции командной строки + +```text +-imacros FILE +-include FILE +``` + +обрабатывают файл до главного input. Они используют обычный preprocessing +engine, а не отдельный облегчённый parser. + +Нормативный порядок начала translation unit: + +```text +predefined macros + -> -D/-U в порядке командной строки + -> все -imacros в порядке командной строки + -> все -include в порядке командной строки + -> главный input +``` + +Таким образом, взаимное расположение `-imacros` и `-include` в `argv` не +перемешивает эти две группы: **все** `-imacros` всегда выполняются раньше +**всех** `-include`. + +`-imacros FILE` полностью обрабатывает файл: его `#define`/`#undef`, +условные директивы, `#lang`/`#endlang`, `#include`, `#include_next`, +`#pragma once` и диагностика имеют обычную семантику. Однако весь normal +preprocessing output этого forced-файла, включая line markers и текст +вложенных headers, отбрасывается. Полученное состояние macro table и других +preprocessing-механизмов сохраняется для последующих forced-файлов и главного +input. + +`-include FILE` использует тот же механизм, но normal output сохраняется, как +если бы найденный header был включён непосредственно перед главным source. +Forced include является настоящей include-границей: внутри него +`__INCLUDE_LEVEL__ == 1`, внутри включённого им header уровень равен `2`, а +главный input остаётся на уровне `0`. `__BASE_FILE__` внутри forced-файлов +остаётся именем главного input. + +Абсолютный operand forced-файла используется непосредственно. Относительный +operand ищется сначала в **current working directory**, затем по обычной +include-chain: + +```text +explicit -I +explicit -isystem +MCPU_CPP_<LANG>_INCLUDE_PATH +MCPU_CPP_INCLUDE_PATH +MCPU_CPP_SYSTEM_INCLUDE_PATH/<lang> +MCPU_CPP_SYSTEM_INCLUDE_PATH +explicit -idirafter +MCPU_CPP_AFTER_INCLUDE_PATH +``` + +Каталог главного input не получает специального приоритета при поиске operand +`-imacros`/`-include`. После нахождения forced-файла обычный quoted +`#include "file"` внутри него снова разрешается относительно физического +каталога этого файла. Если forced-файл найден через элемент include-chain, +его provenance сохраняется и `#include_next` продолжает поиск со следующего +элемента цепочки. + +Forced-файлы и реально достигнутые из них headers входят в обычный physical +dependency registry. Их user/system classification определяется тем же +search provenance, поэтому `-MM`/`-MMD` фильтруют system forced headers так же, +как обычные system headers. Отсутствующий forced-файл является ошибкой. + + +### 4.7. Генерация зависимостей: `-M`, `-MM`, `-MG`, `-MD`, `-MMD`, `-MF`, `-MT`, `-MQ` + +Опции `-M` и `-MM` используют **тот же самый проход include pipeline**, что и +обычная preprocessing. Отдельного повторного поиска заголовков не выполняется. +Поэтому dependency graph автоматически наследует нормативный порядок путей, +`#include_next`, macro-expanded include operands, conditional compilation и +`#pragma once`. + +`-M` подавляет обычный preprocessing output и выводит одно правило Make: + +```make +file.o: file.c header1.h header2.h +``` + +В список входят главный source-файл и все реально достигнутые физические +headers, включая system headers. Один физический файл записывается один раз; +идентичность определяется как `st_dev + st_ino`, поэтому другое относительное +имя, symbolic link или hard link не создают дополнительную dependency. Имя, +назначенное директивой `#line`, является только logical source name и в +dependency list не попадает. + +`-MM` строит тот же граф, но исключает system dependencies. System-контекстом +считаются headers, найденные через explicit `-isystem`, configured system tree +`MCPU_CPP_SYSTEM_INCLUDE_PATH/<lang>` / `MCPU_CPP_SYSTEM_INCLUDE_PATH`, explicit +`-idirafter` и `MCPU_CPP_AFTER_INCLUDE_PATH`, а также вся ветвь headers, +включённая непосредственно или косвенно из такого system header. Форма +`#include "file"` или `#include <file>` сама по себе не определяет system-ness. +Если один и тот же физический файл был достигнут из system-ветви, но затем +также включён непосредственно из user-контекста, он остаётся пользовательской +dependency и присутствует в `-MM`. + +Default target образуется из basename главного source-файла: его suffix +заменяется object suffix (`.o` по умолчанию). Пути и target экранируются для +Make. Для stdin используется GNU-подобная форма `-: -`. + +`-MD` и `-MMD` используют тот же dependency graph, но, в отличие от `-M` и +`-MM`, **не подавляют обычный preprocessing output**. `-MD` включает system +headers, как `-M`; `-MMD` применяет user-only фильтр `-MM`. Это позволяет одним +проходом получить и препроцессированный текст, и side-effect dependency file. + +Если `-MF` не задан, side-effect режим выбирает имя `.d` автоматически: + +* без `-o` из basename входного файла удаляется suffix и добавляется `.d`; + каталоги входного pathname в имя dependency-файла не переносятся; +* при обычном `-o FILE` suffix output-файла заменяется на `.d`; +* для stdin используется имя `-.d`. + +`-MF FILE` переопределяет автоматическое имя dependency-файла. Значение +`-MF -` означает stdout. `-MF` работает также с dependency-only `-M`/`-MM`; +в этом случае оно имеет приоритет над обычным destination для make-rule. Само +по себе `-MF` без одного из `-M`, `-MM`, `-MD`, `-MMD` является ошибкой. + +Семантика намеренно следует GNU CPP: `-MD`/`-MMD` не принимают собственный +аргумент, а `-MF` является отдельной опцией назначения dependency output. + +`-MT TARGET` заменяет автоматический target правила строкой `TARGET` **точно +как она передана**. Make quoting при этом не выполняется. Поэтому один argument +`-MT` может сам содержать несколько targets, разделённых пробелами: + +```text +-MT 'obj/a.o obj/a.pic.o' +``` + +и повторные `-MT` также добавляют targets одного и того же правила: + +```text +-MT obj/a.o -MT obj/a.pic.o +``` + +`-MQ TARGET` имеет ту же семантику выбора target, но экранирует специальные +для Make символы. Например, + +```text +-MQ '$(OBJDIR)/foo.o' +``` + +даёт левую часть правила: + +```make +$$(OBJDIR)/foo.o: +``` + +Поддерживаются как отдельные arguments (`-MT TARGET`, `-MQ TARGET`), так и +attached forms (`-MTTARGET`, `-MQTARGET`). + +Если задан хотя бы один `-MT` или `-MQ`, автоматический default target не +выводится. В частности, `--object-suffix` влияет только на автоматический +target и не переписывает явно заданные targets. Если явных targets нет, +default target экранируется для Make так же, как при `-MQ`. + +Разрешены повторные и смешанные `-MT`/`-MQ`. В соответствии с GNU CPP сначала +выводятся все `-MT` targets в их command-line order, затем все `-MQ` targets в +их command-line order. Все они образуют левую часть **одного** dependency +rule. + +`-MT` и `-MQ` имеют смысл только вместе с одним из dependency-generation +режимов `-M`, `-MM`, `-MD` или `-MMD`. Без такого режима это ошибка командной +строки. + +`-MG` изменяет только обработку **отсутствующих** include-файлов при +dependency-only режимах `-M` и `-MM`. Без `-MG` неразрешённый `#include` остаётся +ошибкой. С `-M -MG` или `-MM -MG` отсутствующий header считается будущим +generated file: preprocessing не завершается ошибкой, а operand директивы +добавляется в dependency rule **ровно в том виде, который получен после macro +expansion**, без приписывания предполагаемого include-directory. Например: + +```text +#include "generated.h" +``` + +при `-M -MG` добавляет dependency `generated.h`, даже если такого файла ещё нет. +Macro-expanded include ведёт себя аналогично: dependency получает уже +развёрнутое имя. `-MG` разрешён только вместе с `-M` или `-MM`; комбинации с +`-MD`/`-MMD` и использование без dependency-only режима являются ошибкой +командной строки. + +Unresolved dependencies интегрированы в **тот же упорядоченный dependency +registry**, что и физически найденные файлы, но образуют отдельный identity +domain. Для найденного файла registry по-прежнему использует `st_dev/st_ino` и +physical provenance. Для отсутствующего файла этих данных нет, поэтому `-MG` +entry не выполняет `stat()` и дедуплицируется по точному тексту include operand. +Это принципиально: наличие одноимённого файла в CWD не должно превращать +неразрешённый `<name>` в ложное physical совпадение, если angle-search этот файл +не находил. Разные unresolved spellings (`generated.h` и `./generated.h`) +считаются разными dependencies. + +Для `-MM` unresolved dependency получает user/system class из контекста поиска: +отсутствующий `<file>` является system-class, отсутствующий `"file"` — user-class, +если сама включающая единица не является system header; любой missing include, +достигнутый из system header, остаётся system-class. При повторении одного и того +же unresolved operand сохраняется классификация его первого появления, что +соответствует GNU CPP. Физически найденные зависимости сохраняют прежнее правило: +если один и тот же inode позднее достигается из user-контекста, он перестаёт быть +system-only. + +`-MG` распространяется также на отсутствующие command-line forced files +`-include FILE` и `-imacros FILE`: их operand заносится в unresolved registry как +user dependency без синтетического search prefix. При наличии реального файла +`-include`/`-imacros` продолжают использовать обычный physical dependency +registry и search provenance. + + +## 5. Переключение языков + +Препроцессор запускается в состоянии `0`. Это безымянный основной C-подобный +язык и он не является допустимым аргументом `#lang`. + +Допустимые языки: + +| Имя | Назначение | +|---|---| +| `diff` | дифференциальные уравнения | +| `dift` | разностные уравнения | +| `alg` | алгебраические уравнения | +| `as` | MCPU assembler (`mcpu-as`) | +| `avm` | схемы аналоговых вычислительных машин | +| `ACS` | структурные схемы систем автоматического управления | + +После `#lang` обязательна строковая константа с одним непустым словом: + +```text +#lang "diff" +``` + +Имя проверяется только по внутреннему списку языков выше и сравнивается без +учёта ASCII-регистра. Поэтому `"diff"`, `"Diff"`, `"DIFF"` и `"dIfF"` +эквивалентны при выборе языка. Исходное написание внутри кавычек при этом +сохраняется в выходном потоке. + +Пробелы внутри строковой константы запрещены: `" diff"`, `"diff "` и +`"di ff"` являются ошибками. Escape-последовательности внутри неё не +разбираются. Закрывающая кавычка обязана находиться на той же физической строке +исходного файла. После неё до конца строки допустимы только пробельные символы. + +Внешние пробелы директивы нормализуются. Например: + +```text + # lang "DiFf" +``` + +превращается в: + +```text +#lang "DiFf" +``` + +`#lang` помещает новый язык в стек, `#endlang` восстанавливает предыдущий. +Стек не сбрасывается при `#include`, поэтому начало и конец языкового блока +могут находиться в разных файлах. Директивы `#lang` и `#endlang` сохраняются в +выходном потоке для последующего frontend dispatcher; `#lang` сохраняется в +нормализованной форме. + +## 6. Простые макроопределения + +Начиная с 0.0.4 поддерживаются object-like macros: + +```text +#define BUFFER_SIZE 1024 +#define NAME value +#define EMPTY +``` + +Директива `#define` сама в выходной поток не попадает. В обычном тексте +идентификатор-макро заменяется его replacement list. Replacement затем снова +просматривается на макроимена, поэтому допускается каскадное раскрытие: + +```text +#define A B +#define B 10 +A +``` + +даёт `10`. + +Во время раскрытия конкретное макро временно блокируется. Поэтому +самоссылочные и взаимно-рекурсивные определения не вызывают бесконечной +рекурсии. + +Макроимена не раскрываются внутри строковых и символьных констант. Для `diff` +апостроф сохраняет специальную языковую семантику и не защищает последующий +текст как C character constant. + +Многострочное определение через backslash-newline поддерживается, поскольку +splice выполняется раньше `#define`. + +### 6.1. `#undef` + +```text +#undef NAME +``` + +удаляет object-like macro. Отмена несуществующего определения не является +ошибкой. + +### 6.2. Вычисляемый `#include` + +Аргумент `#include`, который не начинается непосредственно с `"` или `<`, +сначала проходит macro expansion. Поэтому допустимо: + +```text +#define HEADER <diff/model.h> +#include HEADER +``` + +или + +```text +#define HEADER "local.h" +#include HEADER +``` + +Результат раскрытия обязан иметь форму `"file"` или `<file>`. + +## 7. Макро с аргументами + +Macro engine поддерживает +function-like macros: + +```text +#define идентификатор( список аргументов ) текст +``` + +Открывающая скобка в определении должна идти **непосредственно** +после имени макро. Поэтому + +```text +#define F(X) X +``` + +задаёт макро с аргументом, а + +```text +#define F (X) +``` + +задаёт простое object-like macro со строкой замены `(X)`. + +В месте использования между именем function-like macro и открывающей скобкой +пробельные символы допустимы. Если `(` не следует, идентификатор не считается +вызовом данного макро и остаётся в выходном тексте. + +Для обычного function-like macro число фактических аргументов должно совпадать +с числом формальных. Для variadic macro должны присутствовать все фиксированные +аргументы, а variadic tail может содержать произвольное число аргументов, включая +пустой tail. При разборе списка фактических аргументов вложенные круглые скобки +учитываются; запятая внутри них не разделяет аргументы. Квадратные скобки такого +свойства не имеют — это является частью принятой семантики macro expansion. + +Например, + +```text +#define min(X, Y) ((X) < (Y) ? (X) : (Y)) +min(1, 2) +``` + +даёт + +```text +((1) < (2) ? (1) : (2)) +``` + +Перед подстановкой обычный фактический аргумент сам проходит macro expansion. +Поэтому каскадные и вложенные вызовы работают естественно: + +```text +#define A 7 +#define min(X, Y) ((X) < (Y) ? (X) : (Y)) +min(min(A, 3), 10) +``` + +Формальный параметр может встречаться в replacement list произвольное число +раз. Это означает, что выражение с побочным эффектом в +фактическом аргументе также может быть вычислено несколько раз уже последующим +компилятором; препроцессор не пытается исправлять такую программу. + +Поддерживаются макро без формальных параметров: + +```text +#define READY() 1 +``` + +Они раскрываются только как вызов `READY()` (пробел между именем и `(` при +использовании допустим), но самостоятельный идентификатор `READY` не +раскрывается. + +Имена формальных параметров должны быть различны. Незавершённый список, +неверная пунктуация, недостаточное или избыточное число фактических аргументов +диагностируются как ошибки. + +### 7.1. Stringification `#` + +Поддерживается оператор +stringification (`#`) для параметров function-like macro: + +```text +#define STR(X) #X +STR(alpha + beta) +``` + +даёт + +```text +"alpha + beta" +``` + +Stringification использует **сырой фактический аргумент до macro expansion**. +Поэтому: + +```text +#define A 7 +#define STR(X) #X +#define XSTR(X) STR(X) + +STR(A) -> "A" +XSTR(A) -> "7" +``` + +Ведущие и завершающие пробелы аргумента удаляются. Последовательности +пробельных символов внутри аргумента сворачиваются в один пробел, кроме +пробелов внутри строковых/символьных токенов соответствующего активного +языка. Двойные кавычки и обратные косые черты внутри quoted tokens экранируются +так, чтобы результат оставался одной корректной строковой константой. + +Оператор `#` в replacement list function-like macro обязан непосредственно или +через пробельные символы ссылаться на имя формального параметра. Внутри quoted +token символ `#` оператором не является. Пустой фактический аргумент допустим и +stringify-ится как `""`. + +### 7.2. Token concatenation `##` + +Начиная с 0.0.21 поддерживается оператор token concatenation (`##`) в модели, +согласованной с GNU CPP и механизмом `collect_expansion()` / `macroexpand()` +macro engine. Оператор объединяет два соседних preprocessing token в +один token, после чего получившийся replacement list снова проходит macro +expansion. + +Например: + +```text +#define CAT(A, B) A ## B +CAT(foo, bar) +``` + +даёт `foobar`. Склеивание может образовывать identifier, preprocessing number +или многосимвольный punctuator. Поэтому, например, допустимы: + +```text +CAT(1.5, e3) -> 1.5e3 +CAT(+, =) -> += +``` + +Если формальный параметр непосредственно примыкает к `##`, его фактический +аргумент подставляется **без предварительного macro expansion**. Это тот же +raw-argument принцип, который используется для stringification. Для получения +сначала expansion, а затем concatenation применяется обычный двухуровневый +приём GNU CPP: + +```text +#define AFTERX(X) X_ ## X +#define XAFTERX(X) AFTERX(X) +#define TABLESIZE 1024 +#define BUFSIZE TABLESIZE + +AFTERX(BUFSIZE) -> X_BUFSIZE +XAFTERX(BUFSIZE) -> X_1024 +``` + +Пустой фактический аргумент рядом с `##` ведёт себя как placemarker: сам по +себе он не добавляет token, а `##` с такой стороны не изменяет оставшийся +операнд. Если фактический аргумент содержит несколько preprocessing tokens, +склеивается только крайний token, непосредственно соседний с `##`; остальные +tokens сохраняются и затем участвуют в общем rescan. + +`#` и `##` могут использоваться в одном function-like macro, например: + +```text +#define COMMAND(NAME) #NAME | NAME ## _command +``` + +При этом `#NAME` использует raw spelling аргумента для stringification, а +`NAME ## _command` — тот же raw argument для concatenation. + +`##` внутри quoted token оператором не является. Комментарии к моменту macro +expansion уже заменены whitespace, поэтому они не могут быть созданы +склеиванием `/` и `*`. Между `##` и его операндами исходно может находиться +whitespace; при склеивании он не участвует. + +Если два операнда не образуют один допустимый preprocessing token, выдаётся +диагностика, а сами исходные tokens сохраняются; наличие whitespace между ними +после такой диагностики не является частью контракта. `##` в начале или в +конце replacement list является ошибкой определения macro. + +### 7.3. Variadic macros: `...` и `__VA_ARGS__` + +Начиная с 0.0.46 поддерживаются variadic function-like macros в современной +C99-совместимой форме: + +```text +#define LOG(...) output(__VA_ARGS__) +#define LOGF(format, ...) output(format, __VA_ARGS__) +``` + +Маркер `...` может быть единственным параметром либо последним элементом после +одного или нескольких фиксированных параметров. Старое GNU-расширение с +именованным variadic parameter + +```text +#define LOG(args...) ... +``` + +в 0.0.46 намеренно не поддерживается. `__VA_OPT__` также не является частью +этого релиза. + +При вызове все tokens после последнего фиксированного параметра, включая +разделяющие их запятые, образуют один logical variable argument и подставляются +вместо `__VA_ARGS__`. В обычной позиции этот variable argument предварительно +проходит macro expansion так же, как обычный фактический аргумент: + +```text +#define A 7 +#define V(...) <__VA_ARGS__> +#define F(first, ...) first | __VA_ARGS__ + +V(A, 2, 3) -> <7, 2, 3> +F(1, A, 3) -> 1 | 7, 3 +``` + +Variadic tail может быть пустым. Поэтому оба вызова + +```text +F(1) +F(1,) +``` + +допустимы и подставляют пустой `__VA_ARGS__`. Это **не** означает автоматическое +удаление запятой, явно записанной в replacement list. Например для + +```text +#define E(format, ...) output(format, __VA_ARGS__) +``` + +вызов `E("ok")` оставляет запятую перед пустым tail. Специальная историческая +GNU-семантика `, ## __VA_ARGS__`, удаляющая такую запятую, в контракт 0.0.46 не +входит; если она понадобится, её следует вводить отдельным явно документированным +расширением. + +`__VA_ARGS__` участвует в уже существующей семантике `#` и `##` как настоящий +macro parameter. Stringification использует raw spelling всего variadic tail: + +```text +#define STRV(...) #__VA_ARGS__ +STRV(A, b + c) -> "A, b + c" +``` + +При соседстве с `##` variadic argument также подставляется без prescan; затем +работают обычные правила placemarker, token concatenation и общего rescan. +Например: + +```text +#define L(...) pre ## __VA_ARGS__ +#define R(...) __VA_ARGS__ ## post + +L(fix) -> prefix +R(fix) -> fixpost +``` + +Если variadic argument содержит несколько preprocessing tokens, склеивается +только крайний token, непосредственно соседний с `##`, а остальные tokens +сохраняются, как и для обычного параметра. Пустой tail рядом с `##` ведёт себя +как placemarker. + +Имя `__VA_ARGS__` зарезервировано для variable argument и не принимается как +обычное имя формального параметра. Оператор `#__VA_ARGS__` допустим только в +variadic macro. Dump-режимы сохраняют variadic форму определения, например: + +```text +#define F(first,...) first | __VA_ARGS__ +``` + + +### 7.4. `__VA_OPT__` + +Начиная с 0.0.47 variadic macros поддерживают стандартный условный fragment +`__VA_OPT__(pp-tokens)`. Если variable argument после обычной macro substitution +не содержит preprocessing tokens, весь `__VA_OPT__(...)` раскрывается в пустую +последовательность. Если variable argument непуст, содержимое круглых скобок +участвует в replacement list: + +```text +#define DEBUG(format, ...) \ + fprintf(stderr, format __VA_OPT__(,) __VA_ARGS__) + +DEBUG("ready") -> fprintf(stderr, "ready") +DEBUG("x=%d", x) -> fprintf(stderr, "x=%d", x) +``` + +Решение о непустоте принимается **после expansion variable argument**, а не по +его исходному spelling. Поэтому macro, который сам раскрывается в пустую +последовательность, не активирует `__VA_OPT__`: + +```text +#define EMPTY +#define HAS(...) [__VA_OPT__(yes)] + +HAS() -> [] +HAS(EMPTY) -> [] +HAS(token) -> [yes] +``` + +Содержимое `__VA_OPT__` может включать сбалансированные вложенные круглые скобки. +Закрывающая `)` самого `__VA_OPT__` определяется с учётом их вложенности. +Вложенный `__VA_OPT__` внутри другого `__VA_OPT__` намеренно запрещён. + +`__VA_OPT__` интегрирован с существующими правилами parameter substitution, +stringification, token concatenation, placemarker и rescan. Например: + +```text +#define X 123 +#define S(...) #__VA_OPT__(__VA_ARGS__) +#define L(...) pre ## __VA_OPT__(__VA_ARGS__) + +S() -> "" +S(X) -> "123" +L() -> pre +L(X) -> pre123 +``` + +При `#__VA_OPT__(...)` сначала выполняется parameter substitution внутри +fragment, включая prescan обычных параметров, но произвольные macro names самого +fragment до stringification дополнительно не rescanning-ятся. Поэтому: + +```text +#define X 123 +#define S(a, ...) #__VA_OPT__(a X) + +S(X, y) -> "123 X" +``` + +Если parameter внутри `__VA_OPT__` непосредственно участвует во внутреннем +`##`, для него, как обычно, prescan подавляется; paste выполняется до дальнейшего +rescan. Внешний `##`, соседний с `__VA_OPT__`, получает крайний token уже +подготовленного fragment. Пустой результат `__VA_OPT__` рядом с `##` ведёт себя +как placemarker. + +`__VA_OPT__` допустим только в replacement list variadic function-like macro и +должен непосредственно задавать parenthesized fragment. `##` не может быть +первым или последним preprocessing token внутри самого `__VA_OPT__`. + +Историческое GNU-расширение + +```text +, ## __VA_ARGS__ +``` + +в `mcpu-cpp` намеренно **не реализуется**. Для условной запятой следует +использовать современную форму `__VA_OPT__(,)`. Старое GNU-расширение с +именованным variadic parameter `args...` также остаётся неподдерживаемым. + + + +### 7.5. Нормализация пробелов в replacement list + +Начиная с 0.0.48 `mcpu-cpp` не переносит в результат разворачивания +служебное выравнивание многострочного macro. После удаления `\` + newline +последовательность пробельных символов, принадлежащая самому replacement list, +канонизируется в один ASCII-пробел. Это особенно важно для определений, где +обратные косые черты визуально выровнены в одну колонку: + +```text +#define TRACE(x) \ + do \ + { \ + output(x); \ + done(); \ + } \ + while( 0 ) +``` + +При разворачивании такое определение выдаёт компактный replacement: + +```text +do { output(x); done(); } while( 0 ) +``` + +а не сохраняет десятки пробелов перед каждой бывшей границей физической +строки. + +Нормализация относится **только к whitespace самого replacement list**. +`mcpu-cpp` не является formatter-ом исходной программы: пробелы в обычном +тексте input сохраняются. Пробелы внутри фактического macro argument также не +переформатируются только потому, что argument был подставлен в macro: + +```text +#define ID(x) x + +ID(a + b) -> a + b +``` + +Содержимое string/character literals сохраняется буквально, поэтому: + +```text +#define S "left right" +``` + +по-прежнему содержит пять пробелов внутри строки. + +Наличие whitespace между preprocessing tokens сохраняется как один пробел. +Это не позволяет случайно изменить tokenization, например превратить `+ +` в +`++`, `- >` в `->` или `< <` в `<<`. Операторы `#` и `##`, placemarkers, +`__VA_ARGS__`, `__VA_OPT__` и последующий rescan продолжают использовать свои +существующие правила; новая политика меняет только количество обычного +replacement-list whitespace. + +Dump-режимы (`-dM`, `-dD`) показывают ту же каноническую форму replacement +list, которая хранится во внутренней таблице macro. + + +### 7.6. Компактификация невидимых строк и linemarkers + +Начиная с 0.0.49 `mcpu-cpp` использует для вертикального whitespace ту же +модель, что GNU CPP: **удаляем, но не забываем**. Строки, которые после +preprocessing не породили ни одного выводимого preprocessing token, не обязаны +оставаться физическими пустыми строками в `.E`, однако их исходная позиция +продолжает учитываться при построении linemarkers и значении `__LINE__`. + +Причина невидимости не имеет значения. Это могут быть удалённые directives, +неактивные ветви `#if`, однострочные и многострочные comments, обычные пустые +строки или их смесь. Emitter сравнивает текущую output source position с +позицией следующей реально выдаваемой строки. + +Если следующая позиция находится менее чем через восемь строк, разрыв +представляется обычными newline. Если расстояние равно восьми строкам или +больше, вместо длинной последовательности пустых строк выдаётся корректирующий +linemarker: + +```text +# N "file" +``` + +и следующая содержательная строка сразу относится к source line `N`. Таким +образом граница поведения совместима с GNU CPP: gaps 0..7 сохраняются через +newline, gap 8 и больше заменяется linemarker. + +Structural markers входа и возврата из include-файла сохраняют обычный смысл: + +```text +# 1 "header.h" 1 +# 4 "source.c" 2 +``` + +Если included file не породил никакого output, `mcpu-cpp` не создаёт +искусственный marker, сообщающий, до какой внутренней строки header дошёл +препроцессор. После enter-marker сразу может следовать return-marker. Реальная +позиция снова уточняется только тогда, когда требуется выдать следующий +содержательный текст. + +Эта оптимизация меняет только представление output stream. Source coordinates, +`__LINE__`, diagnostics, `#line`, include enter/return semantics и обработка +macro остаются привязаны к исходному логическому потоку, а не к количеству +физических строк в сжатом `.E`. + + +## 8. Предопределённые макро + +Начиная с 0.0.6 был перенесён исторический механизм predefined macros из +препроцессора. Этот механизм оформлен как отдельный ABI/environment layer +будущего безымянного C-подобного языка. Эти определения не являются +декоративными: их имена и значения должны соответствовать либо семантике GNU +CPP, либо явно документированному MCPU/LibMPU contract. + +### 8.1. Динамические source macros + +Следующие predefined macros вычисляются в точке использования: + +| Макро | Раскрытие | +|---|---| +| `__FILE__` | строковая константа с именем текущего входного файла | +| `__LINE__` | десятичный номер текущей строки | +| `__BASE_FILE__` | строковая константа с именем главного входного файла translation unit | +| `__INCLUDE_LEVEL__` | уровень вложенности `#include`; для главного файла равен `0` | +| `__DATE__` | дата запуска препроцессора в форме `"Mmm dd yyyy"` | +| `__TIME__` | время запуска препроцессора в форме `"hh:mm:ss"` | + +`__DATE__` и `__TIME__` получают один timestamp на весь +translation unit. Специальное раскрытие помещается в output без повторного macro +rescan. + +Эти имена находятся в общей macro table, поэтому `#undef` и последующий +`#define` могут осознанно заменить builtin. + +### 8.2. Версия препроцессора + +Начиная с 0.0.8 standalone preprocessor не определяет GCC-имя `__VERSION__`. +Оно относится к compiler environment, которого для будущего high-level языка +пока нет. Собственная версия `mcpu-cpp` имеет отдельное однозначное имя: + +```text +#define __MCPU_CPP_VERSION__ "1.0.2" +``` + +Значение автоматически берётся из `PACKAGE_VERSION`. Когда появится compiler +frontend/driver, его version contract будет определён отдельно и не будет +смешиваться с версией standalone preprocessor. + +### 8.3. Источники истины ABI + +`mcpu-cpp` собирается только GNU GCC. Во время `configure` проект использует +проверенные приёмы из `LibMPU`/`LibMPUIO` `acsite.m4`: GCC predefined macros +определяют native type sizes, byte/word order и machine-register width, а +установленный `<libmpu.h>` является окончательным источником настроек LibMPU. + +В частности, фиксируются и проверяются: + +```text +MPU_REAL_IO_LIMIT +MPU_MATH_FN_LIMIT +MPU_BYTE_ORDER +MPU_WORD_ORDER +BITS_PER_MACHINE_REGISTER +BITS_PER_UNIT_T +sizeof(__mpu_size_t) +sizeof(__mpu_ptrdiff_t) +``` + +`configure` дополнительно проверяет, что byte order и +`BITS_PER_MACHINE_REGISTER`, записанные в LibMPU, согласованы с GCC target, +которым собирается `mcpu-cpp`. `MPU_WORD_ORDER` берётся непосредственно из +configured LibMPU profile и описывает порядок слов MCPU data environment. + +Пределы `MPU_REAL_IO_LIMIT` и `MPU_MATH_FN_LIMIT` имеют разные назначения. +Например, библиотека может иметь Real I/O до 65536 бит и математические функции +только до 16384 бит. Поэтому `MPU_MATH_FN_LIMIT` не используется как предел +существования типов Real. + +### 8.4. MCPU architecture и assembler prefixes + +Целевая архитектура определяется макро: + +```text +#define _ARCH_MCPU 1 +``` + +MCPU PTR64 имеет ширину 64 бита, поэтому определены `__SIZEOF_POINTER__`, +`__MCPU_POINTER_WIDTH__`, `__INTPTR_TYPE__`, `__UINTPTR_TYPE__`, соответствующие +width/max macros. + +Смысл assembler-prefix macros согласован с GNU CPP, а не с первой буквой имени +register view. В синтаксисе `mcpu-as` дополнительного sigil перед register, +label или immediate нет. `r` и `c` являются частью MCPU register syntax, а не +`REGISTER_PREFIX`. Поэтому: + +```text +#define __REGISTER_PREFIX__ +#define __LOCAL_LABEL_PREFIX__ +#define __USER_LABEL_PREFIX__ +#define __IMMEDIATE_PREFIX__ +``` + +все четыре раскрываются в пустую последовательность. `.L...` остаётся +compiler naming convention и не является assembler ABI local-label prefix: +LOCAL/GLOBAL binding определяется symbol directives. + +### 8.5. Byte order и word order + +Базовые числовые значения порядка байт совместимы с GNU CPP: + +```text +__ORDER_LITTLE_ENDIAN__ +__ORDER_BIG_ENDIAN__ +__ORDER_PDP_ENDIAN__ +``` + +Но целевая среда публикует собственные MCPU names: + +```text +#define __MCPU_BYTE_ORDER__ __ORDER_LITTLE_ENDIAN__ +#define __MCPU_WORD_ORDER__ __ORDER_LITTLE_ENDIAN__ +#define __BYTE_ORDER__ __MCPU_BYTE_ORDER__ +``` + +Фактические значения `__MCPU_BYTE_ORDER__` и `__MCPU_WORD_ORDER__` получают из +configured LibMPU profile (`MPU_BYTE_ORDER` и `MPU_WORD_ORDER`). Поэтому они +следуют host data representation, с которой собрана LibMPU. Это не меняет +отдельный архитектурный контракт кодировки MCPU instruction bytecode. + +GNU/C-specific имя `__FLOAT_WORD_ORDER__` не определяется: типа `float` в +будущем языке MCPU нет. + +Параметры LibMPU/MCPU environment публикуются в MCPU namespace: + +```text +__MCPU_MACHINE_REGISTER_WIDTH__ +__MCPU_REAL_IO_LIMIT__ +__MCPU_MATH_FN_LIMIT__ +__MCPU_INT_MAX_WIDTH__ +__MCPU_REAL_MAX_WIDTH__ +__MCPU_COMPLEX_MAX_WIDTH__ +``` + +`__MCPU_INT_MAX_WIDTH__` равен `NB_I_MAX * 8`, а Real/Complex maximum width +равен configured `MPU_REAL_IO_LIMIT`. `__MCPU_MACHINE_REGISTER_WIDTH__` является +значением `BITS_PER_MACHINE_REGISTER` установленной LibMPU. Пределы Real I/O и +math functions не смешиваются: `MPU_REAL_IO_LIMIT` определяет существование +Real/Complex type family и text conversion, а `MPU_MATH_FN_LIMIT` — наличие +математических функций соответствующей ширины. + +### 8.6. MCPU size/ssize, `ptrdiff` и pointers + +Будущий язык не наследует variable-width C names `short`, `int`, `long` и +не использует C-style имя `size_t` как часть собственного ABI. Беззнаковый +LibMPU size type и знаковый byte-count/error type публикуются симметрично в +MCPU namespace. Например для 64-bit configured profile: + +```text +#define __MCPU_SIZE_TYPE__ uint64 +#define __MCPU_SIZE_WIDTH__ 64 +#define __MCPU_SIZEOF_SIZE__ 8 +#define __MCPU_SIZE_MAX__ 0xffffffffffffffff + +#define __MCPU_SSIZE_TYPE__ int64 +#define __MCPU_SSIZE_WIDTH__ 64 +#define __MCPU_SIZEOF_SSIZE__ 8 +#define __MCPU_SSIZE_MAX__ 0x7fffffffffffffff +``` + +Это MCPU-specific family, а не попытка приписать GNU CPP несуществующий +стандартный `__SSIZE_*` contract. + +MCPU pointer ABI от host не зависит: PTR64 всегда имеет ширину 64 бита: + +```text +#define __INTPTR_TYPE__ int64 +#define __UINTPTR_TYPE__ uint64 +#define __INTPTR_WIDTH__ 64 +#define __UINTPTR_WIDTH__ 64 +#define __INTPTR_MAX__ 0x7fffffffffffffff +#define __UINTPTR_MAX__ 0xffffffffffffffff +#define __SIZEOF_POINTER__ 8 +#define __MCPU_POINTER_WIDTH__ 64 +``` + +Разность MCPU pointers является знаковой и также фиксирована независимо от +host: + +```text +#define __PTRDIFF_TYPE__ int64 +#define __PTRDIFF_WIDTH__ 64 +#define __SIZEOF_PTRDIFF__ 8 +#define __PTRDIFF_MAX__ 0x7fffffffffffffff +``` + +Computed MIN expressions вроде `(-__PTRDIFF_MAX__ - 1)` в predefined table не +создаются. + +### 8.7. Character types + +Обычного C `char` в будущем языке нет. Поэтому `__CHAR_TYPE__` и +`__WCHAR_TYPE__` не определяются. Типы языка называются без C/C++ suffix `_t`: + +```text +#define __CHAR8_TYPE__ char8 +#define __CHAR16_TYPE__ char16 +#define __CHAR8_WIDTH__ 8 +#define __CHAR16_WIDTH__ 16 +#define __SIZEOF_CHAR8__ 1 +#define __SIZEOF_CHAR16__ 2 +``` + +Это типы будущего языка. Внутренняя реализация самого `mcpu-cpp` по-прежнему +использует LibMPUIO `__mpu_char16_t` и strict UCS-2 text model. + +### 8.8. Integer families LibMPU + +Полная structural metadata integer families строится не по жёстко записанному +последнему типу, а до `NB_I_MAX * 8` фактически установленной LibMPU. Для +каждой power-of-two ширины от 8 бит определяются TYPE, WIDTH и SIZEOF: + +```text +#define __INT1024_TYPE__ int1024 +#define __UINT1024_TYPE__ uint1024 +#define __INT1024_WIDTH__ 1024 +#define __UINT1024_WIDTH__ 1024 +#define __SIZEOF_INT1024__ 128 +#define __SIZEOF_UINT1024__ 128 +``` + +На текущей LibMPU 1.0.25 `NB_I_MAX == 8192`, поэтому family доходит до +`int65536`/`uint65536`, а `__SIZEOF_INT65536__ == 8192`. + +Decimal-digit metadata определяется для **каждой** разрешённой integer width: + +```text +__INT<bits>_DECIMAL_DIG__ +__UINT<bits>_DECIMAL_DIG__ +``` + +Значение вычисляется собственными integer-only helpers `mcpu-cpp` из известной +ширины типа. Оно означает точное число десятичных цифр максимального значения +соответствующего типа: знак и завершающий NUL в `DECIMAL_DIG` не входят. Для +unsigned используется максимум `2^bits - 1`, для signed — `2^(bits-1) - 1`. +Это отличается от LibMPU `_int_digs()`, которая предназначена для оценки +строкового буфера и включает место для завершающего NUL. + +Например: + +```text +#define __INT64_DECIMAL_DIG__ 19 +#define __UINT64_DECIMAL_DIG__ 20 +#define __INT256_DECIMAL_DIG__ 77 +#define __UINT256_DECIMAL_DIG__ 78 +``` + +Только сами textual maxima намеренно ограничены шириной `bits <= 256`: + +```text +__INT128_MAX__ +__UINT128_MAX__ +``` + +Максимумы строятся через LibMPU `iuitoa()`. Макро `__INT<bits>_MIN__` не +создаются: predefined table не должна содержать вычисляемые выражения вида +`(-__INT<bits>_MAX__ - 1)`. Для widths больше 256 бит отсутствуют только MAX; +TYPE/WIDTH/SIZEOF/DECIMAL_DIG сохраняются до полного `NB_I_MAX * 8`. + +### 8.9. Real и Complex families LibMPU + +Real/Complex structural metadata генерируется для каждой power-of-two ширины от +32 бит до фактического configured `MPU_REAL_IO_LIMIT`. Для всех этих типов +публикуются TYPE, WIDTH и SIZEOF. + +Для Complex WIDTH означает параметр типа, а не суммарную storage width: + +```text +#define __COMPLEX128_TYPE__ complex128 +#define __COMPLEX128_WIDTH__ 128 +#define __SIZEOF_COMPLEX128__ 32 +``` + +`complex128` состоит из двух компонентов `real128`, поэтому его storage size +равен 32 байтам. При `MPU_REAL_IO_LIMIT == 65536` верх family имеет вид: + +```text +#define __COMPLEX65536_TYPE__ complex65536 +#define __COMPLEX65536_WIDTH__ 65536 +#define __SIZEOF_COMPLEX65536__ 16384 +``` + +Для Real соответственно: + +```text +#define __REAL65536_TYPE__ real65536 +#define __REAL65536_WIDTH__ 65536 +#define __SIZEOF_REAL65536__ 8192 +``` + +Precision metadata определяется для **всех** разрешённых Real widths вплоть +до `MPU_REAL_IO_LIMIT`. Имена macros согласованы с LibMPU helpers: + +```text +__REAL<bits>_DECIMAL_DIG__ -> _real_digs(bits/8) +__REAL<bits>_MANT_DIG__ -> _real_mant_digs(bits/8) +``` + +`__REAL<bits>_DIG__` намеренно отсутствует. Ограничение `bits <= 256` относится +только к большим textual numeric constants. Для размеров до 256 бит также +определяются: + +```text +__REAL<bits>_MAX__ +__REAL<bits>_MIN__ +__REAL<bits>_EPSILON__ +__REAL<bits>_MAX_EXP__ +__REAL<bits>_MIN_EXP__ +__REAL<bits>_MAX_10_EXP__ +__REAL<bits>_MIN_10_EXP__ +``` + +Например, на LibMPU 1.0.25 для `real128` текущий profile даёт значения вида: + +```text +#define __REAL128_EPSILON__ 2.524354896707237777317531409e-29 +#define __REAL128_MAX__ 4.197157432934775384808581951e+323228496 +#define __REAL128_MIN__ 9.530259619551804292864984035e-323228497 +#define __REAL128_MAX_10_EXP__ 323228496 +#define __REAL128_MAX_EXP__ 1073741823 +#define __REAL128_MIN_10_EXP__ -323228524 +#define __REAL128_MIN_EXP__ -1073741822 +``` + +MAX/MIN/EPSILON создаются самой LibMPU и преобразуются через +`real_to_ascii()`. Exponent constants получают значения через LibMPU exponent +helpers и integer conversion. Для widths больше 256 бит эти numeric predefines отсутствуют, но +TYPE/WIDTH/SIZEOF/DECIMAL_DIG/MANT_DIG продолжаются до `MPU_REAL_IO_LIMIT`. + +Для каждого разрешённого Real type вплоть до `MPU_REAL_IO_LIMIT` также +публикуются две компактные характеристики: + +```text +#define __SIZEOF_REAL128_EXP__ 4 +#define __REAL128_MAX_STRLEN__ 60 +``` + +`__SIZEOF_REALxxx_EXP__` непосредственно получает `_sizeof_exp(NB_Rxxx)`. +`__REALxxx_MAX_STRLEN__` получает `_real_max_string(NB_Rxxx)` и означает +максимальное **количество символов** текстового представления, а не количество +байт. Поэтому для zero-terminated строки нужно резервировать не менее +`__REALxxx_MAX_STRLEN__ + 1` элементов: для `char8` это столько же bytes, а для +`char16` физический объём в bytes вдвое больше. Эти два metadata-macro +определяются и для Real widths больше 256, поскольку сами их значения малы. + +### 8.10. Dump macros: `-dM`, `-dMP` + +Опция: + +```text +mcpu-cpp -dM input.c +``` + +печатает только итоговые **непредопределённые** macros в форме `#define ...`. +К этой группе относятся определения из основного файла и включённых headers, а +также определения командной строки `-D`. Предопределённые macros самого +MCPU-CPP в `-dM` не выводятся. Поэтому `-dM` предназначен прежде всего для +короткой инспекции macro-state, созданного пользовательской программой. + +Опция: + +```text +mcpu-cpp -dMP input.c +``` + +добавляет к тому же итоговому состоянию активные predefined macros MCPU-CPP. +Вывод имеет две последовательные группы: сначала все predefined macros, затем +все непредопределённые macros. Внутри каждой группы определения +детерминированно сортируются по имени. Такое разделение удобно системному +разработчику для инспекции preprocessing ABI и архитектурных свойств текущей +MCPU environment, не смешивая их с пользовательскими определениями. + +Принадлежность к группе определяется происхождением macro, а не его именем. +Macro, заданный через `-D` или `#define`, является обычным даже если его имя +похоже на системное. Если predefined macro был удалён через `#undef`, он не +печатается. Если после этого то же имя снова определено пользователем, новое +определение относится к обычной группе и выводится в её части `-dMP`, а также +в `-dM`. Тем самым оба режима показывают именно **итоговый macro-state**. + +Context-dependent `__FILE__`, `__LINE__`, `__DATE__`, `__TIME__`, +`__BASE_FILE__` и `__INCLUDE_LEVEL__` в статическом dump не печатаются. +Статические ABI/architecture predefined macros и вычисляемые static Real +metadata выводятся в `-dMP`. + +Если input file указан, он сначала полностью препроцессируется, после чего +выводится итоговый macro-state; обычный preprocessed text в режимах `-dM` и +`-dMP` не выдаётся. Без input file используется stdin, поэтому пустой stdin с +`-dM` даёт пустой dump, а `-dMP` позволяет получить набор активных static +predefined macros текущей MCPU environment. + +`-dD` имеет другую семантику и этим разделением не затрагивается. + +### 8.11. Dump definitions: `-dD` + +Опция: + +```text +mcpu-cpp -dD input.c +``` + +сохраняет обычный результат препроцессирования и одновременно выводит +встреченные директивы `#define`. Перед началом основного входного текста +печатаются статические предопределённые macro definitions. Каждой такой +дефиниции предшествует marker: + +```text +# 0 "<built-in>" +#define NAME value +``` + +а перед блоком предопределённых macro выводится marker исходного файла вида +`# 0 "input.c"`. Context-dependent `__FILE__`, `__LINE__`, `__DATE__`, +`__TIME__`, `__BASE_FILE__` и `__INCLUDE_LEVEL__` в начальный built-in block +не включаются. + +### 8.12. Dump configuration: `-dconfig` + +Опция: + +```text +mcpu-cpp -dconfig +``` + +не требует input file и выводит effective variables configuration layer после чтения runtime config, необязательного system override, домашнего +user override или выбранного `--config-file`, включая expansion +`$NAME`/`${NAME}`. Строки сортируются по имени и печатаются в форме: + +```text +NAME = value; +``` + +Это позволяет проверить реальные include paths без ручного поиска +`<runtime-root>/etc/mcpu-cpp.conf`, `/etc/mcpu/mcpu-cpp.conf` и +`$HOME/.mcpu/mcpu-cpp.conf`. + +### 8.13. Verbose configuration snapshot: `-v` + +При `-v` MCPU-CPP сохраняет прежний runtime trace для `#lang`, `#include` и +`#include_next`, но конфигурационные переменные печатаются только один раз — +после чтения всех уровней configuration и применения правил приоритета. Поэтому +в verbose output видны только **effective values**, а промежуточные значения из +runtime-root, system и user config не дублируются. + +Config-блок выводится в порядке include policy: language-specific user paths, +общий user path, system root и AFTER path. Переменная, отсутствующая во всех +уровнях configuration, не печатается. Runtime-derived default +`MCPU_CPP_SYSTEM_INCLUDE_PATH` является полноценным самым нижним значением и +поэтому виден при `-v`, даже если ни один `mcpu-cpp.conf` не найден **или все +config-файлы отключены опцией `--no-config`**. + +Форма строки: + +```text +config: NAME=value +``` + +### 8.14. Effective search directories: `-dsearch-dirs` + +Опция: + +```text +mcpu-cpp -dsearch-dirs +``` + +не требует input file, печатает effective глобальные каталоги поиска и +завершает работу без preprocessing. Формат намеренно прост: + +```text +search: /path/to/directory +``` + +Каталоги выводятся в семантическом порядке классов поиска: + +```text +explicit -I +explicit -isystem +configured language-specific user directories +MCPU_CPP_INCLUDE_PATH +MCPU_CPP_SYSTEM_INCLUDE_PATH/<lang> +MCPU_CPP_SYSTEM_INCLUDE_PATH +explicit -idirafter +MCPU_CPP_AFTER_INCLUDE_PATH +``` + +Language-specific entries печатаются для всех поддерживаемых языков в их +каноническом порядке. Во время реального `#include` из этой группы участвует +только каталог активного `#lang`. Каталог текущего физического файла в +`-dsearch-dirs` не выводится: он существует только динамически для конкретного +`#include "..."` и меняется вместе с include stack. `--no-config` не удаляет +runtime-derived system root, поэтому без конфигурационных файлов dump всё равно +содержит `<runtime-root>/include/<lang>` и `<runtime-root>/include`. `-nostdinc` удаляет +из dump effective system `<lang>` entries и system root, но не explicit +`-isystem`. Не существующий на filesystem каталог всё равно показывается, +поскольку он является элементом effective search configuration и просто будет +пропущен при реальном поиске файла. + +`-dsearch-dirs` учитывает `-I`, `-isystem`, `-idirafter`, все уровни config и +replacement-семантику `MCPU_CPP_SYSTEM_INCLUDE_PATH`. Опция `-o` вместе с ним +является ошибкой. + +### 8.15. Условная компиляция + +Директивы `#if`, `#ifdef`, `#ifndef`, `#elif`, `#else` и `#endif` обрабатываются +как управляющие директивы препроцессора и в выходной поток не копируются, в +том числе при `-dD`. Неактивные ветви пропускаются без выполнения находящихся +в них `#define`, `#undef` и `#include`; вложенные условные группы при этом +учитываются корректно. + +Выражение `#if` сначала обрабатывает оператор +`defined`, затем выполняется macro expansion, а оставшиеся идентификаторы +имеют значение `0`. Поддерживаются арифметические, битовые, сравнительные и +логические операции, `?:` и short-circuit semantics для `&&`, `||` и `?:`. + +Начиная с 0.0.26 синтаксис выражения разбирается parser-ом, генерируемым +ZUBR 4.1.0 из `src/mcpp-expr.zubr`; в том же файле находится UCS-2 lexical +analyzer. Предварительная обработка `defined` и macro expansion выполняются до +входа в parser. Арифметическая семантика вынесена в `mcpp-semantic.c/h` и не +зависит от размеров целых типов host-системы. Generated `mcpp-expr.c` +включается в release, поэтому ZUBR требуется только при изменении grammar. + +#### 8.15.1. Единственная вычислительная разрядность — 64 бита + +MCPU-CPP является препроцессором, а не компилятором языка общего назначения. +Все целочисленные вычисления в директивах условной компиляции выполняются +только в 64-разрядной арифметике. Препроцессор не выполняет арифметику LibMPU +произвольной разрядности, вещественные или комплексные вычисления. + +Если программисту не требуется управлять двоичным представлением литерала, +достаточно обычных целых констант и необязательного `U`/`u`. Например: + +```c +#if 2 > 1 +#if 0xffffffffffffffffU > 1 +``` + +Числовой lexeme хранится в UCS-2 до классификации, после чего его ASCII-часть +передаётся LibMPU `iatoui()`. Поддерживаются `0b...`, `0...`, decimal и +`0x...`. Значение, не помещающееся в 64 бита, является ошибкой. Старые C +suffixes `L`, `l`, `LL`, `ll` не поддерживаются. + +#### 8.15.2. Суффикс разрядности `zNNN[Uu]` + +MCPU-CPP понимает общий для MCPU-языков суффикс разрядности: + +```text +zNNN +ZNNN +zNNNu +zNNNU +ZNNNu +ZNNNU +``` + +`NNN` — непустая последовательность десятичных цифр и **всегда** читается как +десятичное число, даже если начинается с нулей. Поэтому `z8`, `z08` и `z008` +задают одну и ту же разрядность 8 бит. + +В общем синтаксисе MCPU корректная разрядность должна быть степенью двойки от +8 до `MPU_REAL_IO_LIMIT`. MCPU-CPP, однако, сознательно ограничен 64-битными +вычислениями: + +* `z8`, `z16`, `z32`, `z64` и варианты регистра допустимы; +* значение `NNN > 64` немедленно является ошибкой: препроцессор не допускает + числовые константы разрядности выше 64 бит в директивах условной компиляции; +* если `NNN <= 64`, но не задаёт допустимую степень двойки, например `z24`, + выводится warning и сам `zNNN` игнорируется; +* необязательный следующий `U`/`u` задаёт unsigned и сохраняет своё значение + даже если некорректный `zNNN` был проигнорирован. + +После полного суффикса должна заканчиваться числовая preprocessing token. +Оператор или punctuation начинает следующий token, поэтому допустимы +`1z32u+2`, `(1z32u)` и `1z32u==1`. Записи вроде `1z32undefined`, `1z32ufoo` и +`1z32$foo` являются ошибками и не разбиваются искусственно на число и имя. + +#### 8.15.3. Нормализация литерала + +Суффикс разрядности действует **только один раз — при формировании значения +самой константы**. Разрядность не сохраняется в semantic value и не участвует +в последующих операциях. + +Для `VALUEzNNN` значение считается знаковым N-битным числом в дополнительном +коде: + +1. сохраняются младшие `NNN` бит; +2. результат расширяется со знаком до 64 бит. + +Для `VALUEzNNNu`/`VALUEzNNNU` сохраняются младшие `NNN` бит, после чего +выполняется нулевое расширение до 64 бит. + +Например: + +```text +0x7fz8 -> 0x000000000000007f -> 127 +0x80z8 -> 0xffffffffffffff80 -> -128 +0xffz8 -> 0xffffffffffffffff -> -1 +0x80z8u -> 0x0000000000000080 -> 128 +0xffz8u -> 0x00000000000000ff -> 255 +0x1ffz8 -> 0xffffffffffffffff -> -1 +0x1ffz8u -> 0x00000000000000ff -> 255 +``` + +Последние два примера намеренны: `zNNN` задаёт разрядность **двоичного +представления**, а не проверку математического диапазона. Биты старше N +отбрасываются до расширения. + +После этой нормализации никакой `z8`, `z16` или `z32` в вычислительной модели +уже не существует. Внутреннее значение содержит только 64-битный битовый +образ и признак signed/unsigned. + +#### 8.15.4. Все последующие операции — 64-битные + +После нормализации все арифметические, побитовые, сравнительные и логические +операции выполняются над 64-битными операндами. Результат операции не +усекается обратно до разрядности исходного suffix. Поэтому: + +```text +0x7fz8 + 1 -> 128 +0xffz8u + 1 -> 256 +``` + +а не `-128` и `0` соответственно. Аналогично `~0xffz8u` инвертирует все 64 +бита и даёт `0xffffffffffffff00`. + +Для бинарных операций, где signedness имеет значение, наличие unsigned +операнда переводит операцию в 64-битную unsigned-интерпретацию. Сравнения +возвращают `0` или `1`. Логические `!`, `&&`, `||` также возвращают signed +64-битные `0` или `1`; short-circuit не вычисляет невыбранную часть. + +Сдвиги выполняются после 64-битной нормализации. Правый сдвиг signed +отрицательного значения является арифметическим, unsigned — логическим. +Например: + +```text +0x80z8 >> 1 -> -64 +0x80z8u >> 1 -> 64 +``` + +Историческое правило MCPU-CPP для отрицательного счётчика сдвига сохраняется: +`A << -N` эквивалентно `A >> N`, а `A >> -N` — `A << N`. + +Таким образом, `zNNN` не превращает препроцессор в компилятор с системой +integer promotions разных размеров. Он лишь позволяет явно описать битовый +образ исходного литерала; затем выражение вычисляется в единственной простой +64-битной модели. + +#### 8.15.5. Символьные константы + +Символьная единица имеет тип `__mpu_uint16_t`, соответствующий внутреннему +UCS-2 представлению, и перед вычислением расширяется нулями до 64 бит. +Последующая арифметика снова является обычной 64-битной арифметикой. + +Состояние условной компиляции хранится в отдельном стеке; условная группа не +может пересекать границу include-файла. + +### 8.16. Диагностические директивы `#error` и `#warning` + +MCPU-CPP поддерживает стандартные диагностические директивы: + +```text +#error сообщение +#warning сообщение +``` + +`#error` выдаёт diagnostic уровня error с текущими логическими именем файла и +номером строки и немедленно завершает preprocessing с ошибкой. `#warning` +выдаёт warning с той же source-location information, после чего preprocessing +продолжается. Поэтому предшествующий `#line` влияет на координаты обеих +диагностик. + +Остаток строки после имени директивы +**не подвергается macro expansion**. Например: + +```c +#define MESSAGE expanded +#warning MESSAGE +``` + +печатает `MESSAGE`, а не `expanded`. Это отличает диагностические директивы от +`#if` и `#line`, где macro expansion является частью соответствующего +контракта. + +Комментарии удаляются на обычной preprocessing phase до обработки директивы. +Начальные и конечные пробелы сообщения удаляются, последовательности пробельных +символов между preprocessing tokens сворачиваются в один пробел. Пробелы внутри +кавычек сохраняются. Например: + +```c +#warning one /* comment */ two +#warning "a b" +``` + +дают сообщения соответственно `one two` и `"a b"`. Unicode-текст проходит +через внутреннее UCS-2 представление и выводится во внешнюю диагностику в UTF-8. + +Обе директивы являются управляющими и никогда не копируются в обычный выходной +поток или в `-dD`. В неактивной ветви `#if` они полностью игнорируются, поэтому +обычная защитная конструкция работает ожидаемо: + +```c +#if 0 +#error this error is inactive +#endif +``` + +### 8.17. Управление предупреждениями: `-Wcomment`, `-Wall`, `-Werror` + +MCPU-CPP разделяет обязательные предупреждения, являющиеся частью уже +зафиксированной preprocessing-семантики, и дополнительные классы предупреждений, +которые включаются пользователем. Управление предупреждениями не изменяет +семантику `-dD`, macro expansion, conditional compilation или include search. + +Опции `-Wcomment` и `-Wcomments` являются полными синонимами и включают два +лексических предупреждения: + +* последовательность `/*`, встретившуюся внутри уже открытого `/* ... */` + комментария; +* backslash-newline внутри `//` комментария, из-за которого однострочный + комментарий физически продолжается на следующую строку. + +По умолчанию этот дополнительный класс выключен. `-Wall` включает все +дополнительные warning classes MCPU-CPP; в версии 0.0.40 таким классом является +`-Wcomment`. Формы `-Wno-comment` и `-Wno-comments` выключают его. Как в GNU +warning model, более специфическая настройка имеет приоритет над групповой +независимо от порядка аргументов. Поэтому обе команды: + +```text +mcpu-cpp -Wall -Wno-comment file.c +mcpu-cpp -Wno-comment -Wall file.c +``` + +оставляют comment warnings выключенными. Между настройками одинаковой +специфичности действует последнее указание, например `-Wno-comment -Wcomment` +включает этот класс. + +`-Werror` не включает никаких новых warning classes. Он повышает до error любое +предупреждение, которое в данном запуске действительно было бы выдано, и такой +запуск завершается неуспешно. Это относится как к дополнительным comment +warnings, так и к уже существующим обязательным предупреждениям MCPU-CPP: + +* активной директиве `#warning`; +* недопустимой, но не превышающей 64 бита ширине `zNNN`; +* переопределению macro другим replacement list; +* результату `##`, не образующему один preprocessing token. + +Например: + +```text +mcpu-cpp -Wcomment -Werror file.c +``` + +превращает найденный comment warning в error. В то же время один `-Werror` без +`-Wcomment`/`-Wall` не заставляет MCPU-CPP искать optional comment warnings. + +`-Wno-error` возвращает обычную severity warning. Для `-Werror` и `-Wno-error`, +имеющих одинаковую специфичность, действует последняя опция командной строки. +Так, `-Werror -Wno-error` оставляет warnings предупреждениями, а +`-Wno-error -Werror` снова повышает их до errors. + +В 0.0.40 намеренно не вводятся `-Werror=<class>`, `-Wno-error=<class>`, +`-Wundef`, `-Wunused-macros`, `-Wtraditional` и другие компиляторные классы. +Warning interface MCPU-CPP остаётся компактным и расширяется только тогда, когда +новый класс действительно нужен самому preprocessing language. + +### 8.18. Идентификаторы UCS-2 + +Начиная с 0.0.22 имена preprocessing identifiers больше не ограничены ASCII. +Внутри `mcpu-cpp` текст уже представлен строгим UCS-2, а классификация символов +выполняется locale-independent функциями LibMPUIO 1.0.4, построенными по Unicode +18.0.0. Первый символ идентификатора должен быть `_` или иметь свойство +`XID_Start`; последующие символы должны быть `_`, `$` или иметь свойство +`XID_Continue`. Символ `$` является расширением `mcpu-cpp`: он разрешён только +после первого символа и не может начинать identifier. Это правило едино для +имён и параметров macro, `#undef`, `#ifdef`/`#ifndef`, `defined`, обычного macro +expansion, `#`/`##`. Имена остаются case-sensitive. Surrogate code units +`U+D800..U+DFFF` не являются допустимыми символами identifiers. + +Например, допустимы: + +```c +#define АНДРЕЙ 1 +#define résumé 2 +#define ΩМЕГА 3 +#define VALUE$OLD 4 +``` + +Например, `VALUE$OLD` допустим, а `$VALUE` недопустим, поскольку `$` не является +identifier-start character. + +Combining marks и не-ASCII decimal digits могут входить в identifier в позициях +`XID_Continue`, но не становятся автоматически допустимыми первыми символами. +Синтаксис числовых констант от этого не меняется: его правила остаются правилами +соответствующего языка, а не Unicode `isdigit`. + +### 8.19. Макросы командной строки `-D` и `-U` + +Начиная с 0.0.23 опции `-D` и `-U` являются полноценными действиями +препроцессора. Поддерживаются формы: + +```text +-DNAME +-DNAME=VALUE +-D'FUNC(a,b)=a+b' +-UNAME +``` + +`-DNAME` эквивалентна `#define NAME 1`; наличие `=` с пустой правой частью +задаёт пустой replacement list. Function-like определения используют тот же +macro engine, что и обычный `#define`, включая параметры, `#`, `##` и +последующий rescanning. `-U` использует тот же identifier contract, что и +`#undef`. Действия `-D`/`-U` выполняются в порядке командной строки после +установки predefined macros. + +Только payload опций `-D` и `-U` интерпретируется как UTF-8 и преобразуется в +строгий UCS-2. Имена файлов, `-I`, другие pathname arguments и остальные +аргументы командной строки остаются исходными byte strings и не подвергаются +Unicode-конвертации. + +Начиная с 0.0.25 символ `$` разрешён внутри имени macro, но не в первой +позиции. При передаче `$` из shell пользователь обязан учитывать правила самого +shell: shell обрабатывает `$` **до запуска `mcpu-cpp`**. Одинарные кавычки уже +полностью защищают `$`, например: + +```sh +mcpu-cpp '-DАНДРЕЙ$_Y=62' input.c +``` + +Без кавычек `$` следует экранировать: + +```sh +mcpu-cpp -DАНДРЕЙ\$_Y=62 input.c +``` + +или использовать двойные кавычки с экранированием: + +```sh +mcpu-cpp -D"АНДРЕЙ\$_Y=62" input.c +``` + +Вариант без защиты: + +```sh +mcpu-cpp -DАНДРЕЙ$_Y=62 input.c +``` + +не передаёт написанное имя буквально: `$...` сначала раскрывается shell и +`mcpu-cpp` получает уже изменённый `argv`. Внутри одинарных кавычек обратная +косая черта перед `$` не нужна и стала бы обычным символом аргумента. + +Для command-line `-D` левая часть до первого `=` разбирается как отдельный +macro declarator. Если после допустимого имени (или завершённого списка +параметров function-like macro) до `=` встречается недопустимый хвост, этот +хвост молча отбрасывается и **никогда не превращается в replacement list**. +Например: + +```text +-D'АНДРЕЙ@XYZ=62' +``` + +эквивалентно: + +```c +#define АНДРЕЙ 62 +``` + +а не ошибочной форме `#define АНДРЕЙ @XYZ 62`. Аналогичное правило допустимого +identifier-prefix применяется к `-U`. Если же первый символ вообще не является +допустимым identifier-start character (например `$` или цифра), определение +остаётся ошибочным. + +При `-dD` определения, пришедшие через `-D`, маркируются отдельно от +предопределённых macro: + +```text +# 0 "<command-line>" +#define NAME value +``` + +в то время как predefined macros продолжают использовать `<built-in>`. + +### 8.20. Публичный интерфейс командной строки + +`mcpu-cpp` поддерживает только актуальные опции, описанные `--help`. Устаревшие +compatibility-флаги не образуют скрытый интерфейс и диагностируются как +`unknown option`. Опция `-E` является исключением: она молча принимается и +игнорируется, поскольку может передаваться compiler driver при запуске +отдельного препроцессора. + +Опция `--object-suffix SUFFIX` задаёт суффикс object target, используемый при +генерации make-зависимостей; аргумент обязателен. + +## 9. Build-system и генераторы + +Собственные Autoconf-макросы проекта находятся в корневом `acsite.m4`. +Каталог `m4/` зарезервирован для внешних/vendor M4-файлов. Такой порядок +повторяет принятую в библиотеках MCPU схему и не смешивает собственный +configure-код с импортированными макросами. + +Парсер выражений `#if` генерируется ZUBR 4.1.0 из `src/mcpp-expr.zubr`. +Release archive содержит и грамматику, и уже сгенерированный `src/mcpp-expr.c`, +поэтому обычная сборка не требует установленного ZUBR. После изменения +грамматики developer build использует штатное правило Automake: + +```text +zubr -vl -s -Bmcpp_ -o mcpp-expr.c mcpp-expr.zubr +``` + +Перед выпуском release generated C должен соответствовать грамматике, полный +test suite и `make distcheck` должны проходить без ошибок. + +### 9.1. Developer bootstrap и Git source tree + +Начиная с 0.0.50 корневой скрипт `./bootstrap` позволяет не хранить в Git файлы, +которые полностью воспроизводятся из исходников. Скрипт сначала генерирует +`src/mcpp-expr.c` из `src/mcpp-expr.zubr` с помощью ZUBR 4.1.0, затем выполняет +`aclocal`, `autoheader`, `automake` и `autoconf` в стиле библиотек LibMPU и +LibMPUIO. Опция `--target-dest-dir=DIR` задаёт target ROOTFS для системных +Autoconf macro/include directories. + +Это правило относится именно к developer Git tree. **Release archive остаётся +самодостаточным**, как и раньше: он содержит `configure`, `Makefile.in`, helper +scripts Automake и уже сгенерированный `src/mcpp-expr.c`, поэтому обычная сборка +релиза не требует предварительного запуска `bootstrap` и не требует ZUBR. + +Корневой `.gitignore` перечисляет воспроизводимые bootstrap-файлы и обычный +configure/build state. Он не меняет существующую release/build model, а только +позволяет поддерживать более чистый Git repository. + +## 10. GNU-compatible features + +`mcpu-cpp` является самостоятельным препроцессором MCPU, но для ряда хорошо +известных операций намеренно повторяет поведение GNU CPP. Совместимость +относится к документированным возможностям, а не означает полную CLI- или +языковую взаимозаменяемость с GCC. + +В частности, GNU-compatible поведение используется для: + +* object-like и function-like macro, повторного macro rescan, `#` и `##`; +* variadic macro `...` / `__VA_ARGS__` и стандартного `__VA_OPT__`; +* `#if`, `#ifdef`, `#ifndef`, `#elif`, `#else`, `#endif` и `defined`; +* `#include`, `#include_next`, `#pragma once`, `#line` и GNU linemarkers; +* compact output mapping: до семи невидимых строк представляются newline, а + разрыв в восемь и более строк — корректирующим linemarker; +* forced files `-include` / `-imacros` и dependency options `-M`, `-MM`, `-MD`, + `-MMD`, `-MF`, `-MT`, `-MQ`, `-MG`; +* warning controls `-w`, `-Wall`, `-Werror` и поддерживаемых `-Wcomment` forms. + +MCPU-specific возможности, включая `#lang` / `#endlang`, числовой суффикс +`zNNN` и ABI predefined macros, остаются собственными расширениями `mcpu-cpp`. |
