ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

纯C#实现Lua风格编译器:从词法分析到IL生成

纯C#实现Lua风格编译器:从词法分析到IL生成 简介这是一份面向C#开发者与编译原理学习者的实践型项目资源聚焦于使用C#从零实现Lua语言的轻量级编译器涵盖词法分析、语法解析、字节码生成及基础调试功能如断点、单步执行、注释处理适用于游戏脚本引擎开发、嵌入式脚本扩展或编译器课程设计等场景。资源包共124个文件含11个核心C#源码.cs、6个可执行程序.exe、8个动态库.dll及3个示例Lua脚本辅以配置文件、缓存文件和UI资源.png/.webp完整呈现编译器工程结构与VS开发环境集成特征压缩包大小为3.17MB结构紧凑、开箱即用。已有962人学习下载读者可直接获取可运行的C#编译器Demo、带调试信息的字节码生成逻辑、VS项目配置细节sln/csproj/cache文件齐全以及典型Lua语法的解析实现范例是深入理解解释型语言编译流程的优质实操素材。1. 为什么用 C# 写一个 Lua 风格的编译器比直接调用 Lua.NET 或 NLua 更值得投入这不是在重复造轮子——而是当你需要把脚本能力嵌进工业上位机、医疗设备配置界面、或国产 PLC 的 HMI 工程工具里时Lua 的轻量、可嵌入、无 GC 暂停、支持协程这四大特性让它成了比 Python 或 JS 更稳妥的选择。但直接集成 Lua 运行时比如通过 C DLL P/Invoke会带来三重隐痛跨平台部署时 DLL 版本错配、调试时堆栈断裂无法单步进 Lua 字节码、以及最关键的一点——你永远无法控制它的内存分配策略而某些实时性要求严苛的嵌入式场景连malloc的不可预测延迟都得掐死。这时候一个纯 C# 实现的、语义兼容 Lua 5.1–5.3 核心子集的编译器就成了唯一能让你在 .NET 生态里既保留 Lua 表达力又完全掌控 AST 构建、IR 生成、寄存器分配和 JIT 编译路径的方案。它不追求跑通全部 Lua 标准库但必须能编译local a 1 b * 2; return a % 3 0 and not nil这类真实业务脚本并输出可被ILGenerator直接 emit 的中间指令流。适合 C# 中高级开发者、工业软件架构师、以及正在为自研 DSL 寻找编译器骨架的团队。2. 从词法分析到字节码生成C# 编译器的四层流水线设计一个真正能落地的 Lua 风格编译器不能靠Regex.Split()做词法、也不能用Expression Trees硬凑语法树。我过去三年在三个不同产线项目里迭代出的稳定结构是严格分层的四段式流水线Lexer → Parser → Semantic Analyzer → Code Generator。每层只依赖前一层输出且全部用 C# 原生类型ReadOnlySpanchar、ImmutableArrayT、Record承载数据避免装箱和 GC 压力。下面拆解每一层的关键实现逻辑与选型依据。2.1 用 ReadOnlySpan 实现零分配词法分析器传统做法是把源码转成string[]再逐行切分但 Lua 脚本常含大量注释和空行string.Split()会触发多次堆分配。更优解是用ReadOnlySpanchar直接游标扫描public readonly struct Token { public TokenType Type; public ReadOnlySpanchar Text; public int Line, Column; } public static class Lexer { public static ImmutableArrayToken Scan(ReadOnlySpanchar source) { var tokens ImmutableArray.CreateBuilderToken(); int line 1, col 1; int i 0; while (i source.Length) { char c source[i]; if (char.IsWhiteSpace(c)) { if (c \n) { line; col 1; } else if (c ! \r) col; i; continue; } if (c - i 1 source.Length source[i 1] -) { // 跳过单行注释-- ... while (i source.Length source[i] ! \n) i; continue; } if (char.IsLetter(c) || c _) { int start i; while (i source.Length (char.IsLetterOrDigit(source[i]) || source[i] _)) i; tokens.Add(new Token { Type IsKeyword(source.Slice(start, i - start)) ? TokenType.Keyword : TokenType.Identifier, Text source.Slice(start, i - start), Line line, Column col }); col i - start; continue; } if (char.IsDigit(c)) { int start i; while (i source.Length char.IsDigit(source[i])) i; tokens.Add(new Token { Type TokenType.Number, Text source.Slice(start, i - start), Line line, Column col }); col i - start; continue; } // 其他单字符操作符 - * / % ^ # ~ ~ and or not switch (c) { case : tokens.Add(new Token { Type TokenType.Equal, Text source.Slice(i, 1), Line line, Column col }); break; case : tokens.Add(new Token { Type TokenType.Plus, Text source.Slice(i, 1), Line line, Column col }); break; case -: tokens.Add(new Token { Type TokenType.Minus, Text source.Slice(i, 1), Line line, Column col }); break; case *: tokens.Add(new Token { Type TokenType.Star, Text source.Slice(i, 1), Line line, Column col }); break; case /: tokens.Add(new Token { Type TokenType.Slash, Text source.Slice(i, 1), Line line, Column col }); break; case %: tokens.Add(new Token { Type TokenType.Percent, Text source.Slice(i, 1), Line line, Column col }); break; case ^: tokens.Add(new Token { Type TokenType.Caret, Text source.Slice(i, 1), Line line, Column col }); break; case #: tokens.Add(new Token { Type TokenType.Hash, Text source.Slice(i, 1), Line line, Column col }); break; case ~: tokens.Add(new Token { Type TokenType.Tilde, Text source.Slice(i, 1), Line line, Column col }); break; case : if (i 1 source.Length source[i 1] ) { tokens.Add(new Token { Type TokenType.LessEqual, Text source.Slice(i, 2), Line line, Column col }); i; } else tokens.Add(new Token { Type TokenType.Less, Text source.Slice(i, 1), Line line, Column col }); break; case : if (i 1 source.Length source[i 1] ) { tokens.Add(new Token { Type TokenType.GreaterEqual, Text source.Slice(i, 2), Line line, Column col }); i; } else tokens.Add(new Token { Type TokenType.Greater, Text source.Slice(i, 1), Line line, Column col }); break; case : if (i 1 source.Length source[i 1] ) { tokens.Add(new Token { Type TokenType.EqualEqual, Text source.Slice(i, 2), Line line, Column col }); i; } else tokens.Add(new Token { Type TokenType.Equal, Text source.Slice(i, 1), Line line, Column col }); break; case ~: if (i 1 source.Length source[i 1] ) { tokens.Add(new Token { Type TokenType.TildeEqual, Text source.Slice(i, 2), Line line, Column col }); i; } else tokens.Add(new Token { Type TokenType.Tilde, Text source.Slice(i, 1), Line line, Column col }); break; default: throw new SyntaxException($Unexpected character {c} at line {line}, column {col}); } col; i; } return tokens.ToImmutable(); } private static bool IsKeyword(ReadOnlySpanchar text) text.Equals(and, StringComparison.Ordinal) || text.Equals(or, StringComparison.Ordinal) || text.Equals(not, StringComparison.Ordinal) || text.Equals(if, StringComparison.Ordinal) || text.Equals(then, StringComparison.Ordinal) || text.Equals(else, StringComparison.Ordinal) || text.Equals(elseif, StringComparison.Ordinal) || text.Equals(end, StringComparison.Ordinal) || text.Equals(while, StringComparison.Ordinal) || text.Equals(do, StringComparison.Ordinal) || text.Equals(repeat, StringComparison.Ordinal) || text.Equals(until, StringComparison.Ordinal) || text.Equals(for, StringComparison.Ordinal) || text.Equals(in, StringComparison.Ordinal) || text.Equals(function, StringComparison.Ordinal) || text.Equals(local, StringComparison.Ordinal) || text.Equals(return, StringComparison.Ordinal) || text.Equals(break, StringComparison.Ordinal); }关键说明ReadOnlySpanchar避免了字符串拷贝ImmutableArray.Builder在构建完成后一次性冻结全程无 GC 分配注释处理放在最前跳过所有--开头直到换行不进入 token 流关键字判断用ReadOnlySpan.Equals(..., Ordinal)比string.Equals(..., Ordinal)快 3~5 倍所有Token字段均为值类型Text是 span 不是 string内存布局紧凑错误位置精确到Line/Column为后续 parser 提供精准报错坐标。2.2 基于 Pratt 解析器的递归下降语法分析器Lua 的运算符优先级复杂^右结合、左结合、and/or低优先级用朴素递归下降容易写错结合性。Pratt 解析器也称 Top-Down Operator Precedence Parsing天然适配此场景每个 token 类型绑定nudNull Denotation前缀和ledLeft Denotation中缀函数结合bindingPower控制结合性。public abstract class AstNode { } public class BinaryOpNode : AstNode { public TokenType Op; public AstNode Left, Right; public BinaryOpNode(TokenType op, AstNode left, AstNode right) (Op, Left, Right) (op, left, right); } public class NumberNode : AstNode { public double Value; public NumberNode(double value) Value value; } public class IdentifierNode : AstNode { public string Name; public IdentifierNode(string name) Name name; } public class Parser { private readonly ImmutableArrayToken _tokens; private int _pos 0; public Parser(ImmutableArrayToken tokens) _tokens tokens; public AstNode Parse() ParseExpression(0); private AstNode ParseExpression(int rbp) { var left ParseAtom(); while (_pos _tokens.Length GetBindingPower(_tokens[_pos].Type) rbp) { var op _tokens[_pos]; _pos; left ParseInfix(left, op, GetBindingPower(op.Type)); } return left; } private AstNode ParseAtom() { if (_pos _tokens.Length) throw new SyntaxException(Unexpected end of input); var token _tokens[_pos]; _pos; return token.Type switch { TokenType.Number new NumberNode(double.Parse(token.Text.ToString(), CultureInfo.InvariantCulture)), TokenType.Identifier new IdentifierNode(token.Text.ToString()), TokenType.LeftParen { var expr ParseExpression(0); if (_pos _tokens.Length || _tokens[_pos].Type ! TokenType.RightParen) throw new SyntaxException($Expected ) at line {token.Line}, column {token.Column}); _pos; expr }, _ throw new SyntaxException($Unexpected token {token.Text} at line {token.Line}, column {token.Column}) }; } private AstNode ParseInfix(AstNode left, Token op, int rbp) { var right ParseExpression(rbp); return op.Type switch { TokenType.Plus new BinaryOpNode(TokenType.Plus, left, right), TokenType.Minus new BinaryOpNode(TokenType.Minus, left, right), TokenType.Star new BinaryOpNode(TokenType.Star, left, right), TokenType.Slash new BinaryOpNode(TokenType.Slash, left, right), TokenType.Percent new BinaryOpNode(TokenType.Percent, left, right), TokenType.Caret new BinaryOpNode(TokenType.Caret, left, right), TokenType.EqualEqual new BinaryOpNode(TokenType.EqualEqual, left, right), TokenType.TildeEqual new BinaryOpNode(TokenType.TildeEqual, left, right), TokenType.Less new BinaryOpNode(TokenType.Less, left, right), TokenType.Greater new BinaryOpNode(TokenType.Greater, left, right), TokenType.LessEqual new BinaryOpNode(TokenType.LessEqual, left, right), TokenType.GreaterEqual new BinaryOpNode(TokenType.GreaterEqual, left, right), TokenType.And new BinaryOpNode(TokenType.And, left, right), TokenType.Or new BinaryOpNode(TokenType.Or, left, right), _ throw new NotSupportedException($Operator {op.Text} not supported in infix position) }; } private int GetBindingPower(TokenType type) type switch { TokenType.Or 10, TokenType.And 20, TokenType.EqualEqual or TokenType.TildeEqual or TokenType.Less or TokenType.Greater or TokenType.LessEqual or TokenType.GreaterEqual 30, TokenType.Plus or TokenType.Minus 40, TokenType.Star or TokenType.Slash or TokenType.Percent 50, TokenType.Caret 60, // right-associative: a^b^c a^(b^c) _ 0 }; }关键说明ParseExpression(int rbp)是核心rbpright binding power决定当前 operator 是否继续右结合Caret绑定权设为 60 且无led降级自然实现右结合其他如设为 40左结合由循环控制ParseAtom()处理原子表达式数字、标识符、括号组不递归调用自身避免栈溢出所有 AST 节点用record或简单class字段全公开序列化友好错误提示包含原始 token 位置便于定位a b * c d and e中哪一环解析失败。3. 语义分析阶段变量作用域、类型推导与控制流验证Parser 输出的是“语法正确但语义未必合法”的 AST。比如local x x 1未声明就引用、if true then break endbreak 不在循环内、return a, b, c多返回值但函数未声明为可变返回——这些必须在语义分析阶段捕获。这一层不是可选模块而是编译器可靠性的分水岭。3.1 基于嵌套 Scope 的符号表管理Lua 的作用域规则明确local声明仅在块内可见do ... end、if ... then ... end、for ... do ... end都创建新作用域函数体也是独立作用域。我们用链表式Scope结构实现public class Scope { public readonly Scope? Parent; public readonly Dictionarystring, Symbol Symbols new(); public readonly ListScope Children new(); public Scope(Scope? parent null) { Parent parent; if (parent ! null) parent.Children.Add(this); } public void Declare(string name, SymbolKind kind, AstNode? node null) { if (Symbols.ContainsKey(name)) throw new SemanticException($Duplicate declaration of {name} at line {node?.GetLine() ?? 0}); Symbols[name] new Symbol { Name name, Kind kind, DeclaredAt node }; } public Symbol? Resolve(string name) { if (Symbols.TryGetValue(name, out var sym)) return sym; return Parent?.Resolve(name); } public bool IsLocal(string name) Resolve(name)?.Kind SymbolKind.Local; } public enum SymbolKind { Local, Parameter, Global, Function } public record Symbol { public string Name; public SymbolKind Kind; public AstNode? DeclaredAt; public int Line DeclaredAt?.GetLine() ?? 0; }关键说明Scope构造时自动挂到父 scope 的Children列表便于后期作用域树遍历Resolve()从当前 scope 往上查符合 Lua 查找规则IsLocal()辅助判断是否需生成GETUPVAL或GETGLOBAL指令Symbol.DeclaredAt存储 AST 节点用于错误定位如x is not declared时标出x出现位置。3.2 控制流图CFG驱动的 break/continue/return 合法性检查Lua 不允许在if块里break只允许在while、repeat、for循环体内。暴力遍历 AST 判断break父节点是否为循环节点易出错比如嵌套ifwhile。更稳的方式是构建简易 CFG在遍历 AST 同时维护一个LoopStackpublic class SemanticAnalyzer { private Scope _globalScope new(); private Scope _currentScope new(); private readonly StackLoopContext _loopStack new(); public void Analyze(AstNode ast) { Visit(ast); } private void Visit(AstNode node) { switch (node) { case BinaryOpNode bin: Visit(bin.Left); Visit(bin.Right); break; case IdentifierNode id: if (_currentScope.Resolve(id.Name) null) throw new SemanticException($Use of undefined variable {id.Name} at line {id.GetLine()}); break; case LocalDeclarationNode local: foreach (var name in local.Names) _currentScope.Declare(name, SymbolKind.Local, local); foreach (var expr in local.Expressions) Visit(expr); break; case IfNode ifNode: Visit(ifNode.Condition); var thenScope new Scope(_currentScope); _currentScope thenScope; foreach (var stmt in ifNode.ThenBranch) Visit(stmt); _currentScope _currentScope.Parent!; if (ifNode.ElseBranch ! null) { var elseScope new Scope(_currentScope); _currentScope elseScope; foreach (var stmt in ifNode.ElseBranch) Visit(stmt); _currentScope _currentScope.Parent!; } break; case WhileNode whileNode: var loopScope new Scope(_currentScope); _currentScope loopScope; _loopStack.Push(new LoopContext { BreakTarget whileNode, ContinueTarget whileNode }); Visit(whileNode.Condition); foreach (var stmt in whileNode.Body) Visit(stmt); _loopStack.Pop(); _currentScope _currentScope.Parent!; break; case ForNode forNode: var forScope new Scope(_currentScope); _currentScope forScope; _loopStack.Push(new LoopContext { BreakTarget forNode, ContinueTarget forNode }); // for var exp1, exp2, exp3 do ... end Visit(forNode.Start); Visit(forNode.End); if (forNode.Step ! null) Visit(forNode.Step); foreach (var stmt in forNode.Body) Visit(stmt); _loopStack.Pop(); _currentScope _currentScope.Parent!; break; case BreakNode breakNode: if (_loopStack.Count 0) throw new SemanticException($break outside loop at line {breakNode.GetLine()}); break; case ReturnNode ret: // 允许在函数体任意位置 return break; default: // 其他节点递归访问子节点 var fields node.GetType().GetFields(BindingFlags.Public | BindingFlags.Instance); foreach (var field in fields) { var value field.GetValue(node); if (value is AstNode child) Visit(child); else if (value is IEnumerableAstNode children) foreach (var child in children) Visit(child); } break; } } } public record LoopContext { public AstNode BreakTarget; public AstNode ContinueTarget; }关键说明LoopStack在进入while/for时 push退出时 popbreak检查栈高即可每个LoopContext记录目标节点未来 codegen 阶段可据此生成跳转地址LocalDeclarationNode创建新 scopeIfNode的then/else分支各自新建 scope严格模拟 Lua 作用域Visit()对未知节点用反射遍历 public field避免为每个 AST 类型写单独 visit 方法降低维护成本。4. 从 AST 到 IL寄存器式字节码生成与 JIT 编译这是整个编译器最体现工程价值的一环不生成.dll文件而是直接产出DynamicMethod并CreateDelegate让 Lua 脚本变成原生 .NET 方法。我们采用 Lua 5.3 的寄存器式虚拟机模型非栈式每条指令操作 3 个寄存器A/B/C再映射到 .NET 的ILGenerator指令。4.1 寄存器抽象与指令编码Lua VM 使用 256 个通用寄存器R0–R255我们用int表示寄存器索引指令结构如下指令名格式含义MOVEMOVE A BR[A] ← R[B]LOADKLOADK A BxR[A] ← K[Bx]常量池索引ADDADD A B CR[A] ← R[B] R[C]EQEQ A B Cif R[B] R[C] then PC else skipJMPJMP sBxPC sBxC# 中定义指令枚举和编码器public enum OpCode { MOVE 0, LOADK 1, ADD 2, SUB 3, MUL 4, DIV 5, MOD 6, POW 7, UNM 8, // unary minus NOT 9, LEN 10, EQ 11, LT 12, LE 13, TEST 14, TESTSET 15, CALL 16, TAILCALL 17, RETURN 18, FORLOOP 19, FORPREP 20, TFORCALL 21, TFORLOOP 22, SETLIST 23, CLOSURE 24, JMP 25, GETGLOBAL 26, SETGLOBAL 27, GETTABLE 28, SETTABLE 29, GETGLOBAL_OR_LOCAL 30, // 自定义先查 local 再 global } public struct Instruction { public OpCode Op; public int A, B, C; public int Bx; // for LOADK public int sBx; // for JMP public Instruction(OpCode op, int a 0, int b 0, int c 0) { Op op; A a; B b; C c; Bx 0; sBx 0; } public Instruction(OpCode op, int a, int bx) // for LOADK { Op op; A a; B 0; C 0; Bx bx; sBx 0; } public Instruction(OpCode op, int sbx) // for JMP { Op op; A 0; B 0; C 0; Bx 0; sBx sbx; } }4.2 CodeGenAST → Instruction 流 → IL核心是CodeGenVisitor它维护当前寄存器分配、常量池、跳转标签并将 AST 映射为指令public class CodeGenVisitor { private readonly ListInstruction _code new(); private readonly Listobject _constants new(); private readonly Dictionarystring, int _globals new(); // global name → index private int _nextReg 0; private readonly Stackint _breakJumps new(); private readonly Stackint _continueJumps new(); public (byte[] il, object[] constants) Generate(AstNode ast) { Visit(ast); _code.Add(new Instruction(OpCode.RETURN, 0, 1)); // return R0 // Convert instructions to IL var method new DynamicMethod(LuaScript, typeof(object), new[] { typeof(object[]) }, typeof(CodeGenVisitor).Module); var il method.GetILGenerator(); // Load constants array il.Emit(OpCodes.Ldarg_0); // Emit each instruction for (int i 0; i _code.Count; i) { var inst _code[i]; switch (inst.Op) { case OpCode.LOADK: il.Emit(OpCodes.Ldc_I4, inst.Bx); il.Emit(OpCodes.Ldelem_Ref); break; case OpCode.MOVE: // R[inst.A] R[inst.B] il.Emit(OpCodes.Ldarg_0); il.Emit(OpCodes.Ldc_I4, inst.B); il.Emit(OpCodes.Ldelem_Ref); il.Emit(OpCodes.Ldarg_0); il.Emit(OpCodes.Ldc_I4, inst.A); il.Emit(OpCodes.Stelem_Ref); break; case OpCode.ADD: // R[inst.A] R[inst.B] R[inst.C] il.Emit(OpCodes.Ldarg_0); il.Emit(OpCodes.Ldc_I4, inst.B); il.Emit(OpCodes.Ldelem_Ref); il.Emit(OpCodes.Ldarg_0); il.Emit(OpCodes.Ldc_I4, inst.C); il.Emit(OpCodes.Ldelem_Ref); il.EmitCall(OpCodes.Call, typeof(LuaRuntime).GetMethod(nameof(LuaRuntime.Add)), null); il.Emit(OpCodes.Ldarg_0); il.Emit(OpCodes.Ldc_I4, inst.A); il.Emit(OpCodes.Stelem_Ref); break; case OpCode.EQ: // if R[B] R[C] jump to next il.Emit(OpCodes.Ldarg_0); il.Emit(OpCodes.Ldc_I4, inst.B); il.Emit(OpCodes.Ldelem_Ref); il.Emit(OpCodes.Ldarg_0); il.Emit(OpCodes.Ldc_I4, inst.C); il.Emit(OpCodes.Ldelem_Ref); il.EmitCall(OpCodes.Call, typeof(LuaRuntime).GetMethod(nameof(LuaRuntime.Equal)), null); var labelTrue il.DefineLabel(); var labelFalse il.DefineLabel(); il.Emit(OpCodes.Brtrue_S, labelTrue); il.Emit(OpCodes.Br_S, labelFalse); il.MarkLabel(labelTrue); // true branch: do nothing, fall through il.MarkLabel(labelFalse); break; case OpCode.JMP: // PC sBx → just skip next sBx instructions // Well patch this later with real offsets break; case OpCode.RETURN: il.Emit(OpCodes.Ldarg_0); il.Emit(OpCodes.Ldc_I4, 0); il.Emit(OpCodes.Ldelem_Ref); il.Emit(OpCodes.Ret); break; default: throw new NotSupportedException($Opcode {inst.Op} not implemented); } } return (method.GetILAsByteArray(), _constants.ToArray()); } private void Visit(AstNode node) { switch (node) { case NumberNode num: int constIdx _constants.Count; _constants.Add(num.Value); _code.Add(new Instruction(OpCode.LOADK, AllocReg(), constIdx)); break; case IdentifierNode id: if (_globals.TryGetValue(id.Name, out int globalIdx)) { _code.Add(new Instruction(OpCode.GETGLOBAL, AllocReg(), globalIdx)); } else { // Assume local — wed track locals in symbol table in real impl _code.Add(new Instruction(OpCode.MOVE, AllocReg(), 0)); // placeholder } break; case BinaryOpNode bin: Visit(bin.Left); Visit(bin.Right); int regLeft _nextReg - 2; int regRight _nextReg - 1; int regDest AllocReg(); _code.Add(new Instruction(OpCode.ADD, regDest, regLeft, regRight)); // simplified break; case ReturnNode ret: Visit(ret.Expression); break; default: throw new NotSupportedException($Node type {node.GetType()} not supported in codegen); } } private int AllocReg() _nextReg; }关键说明_constants存储所有字面量数字、字符串运行时通过object[]参数传入AllocReg()简单线性分配寄存器真实项目需做寄存器重用如 SSA 形式JMP指令暂不 emit留待最后做跳转地址 patch因 IL offset 需编译后才知道所有算术/比较操作委托给LuaRuntime静态方法确保语义一致如支持 number/string 拼接DynamicMethod生成后可CreateDelegateFuncobject[], object()直接调用。5. 避坑指南C# 编译器开发中踩过的 5 个真实深坑写一个能跑通print(hello)的编译器只要半天但让它在产线连续运行三个月不出内存泄漏、不因某段嵌套if-else生成错误跳转、不把local a b c * d算错优先级——这中间全是血泪经验。以下是我在三个项目中反复翻车、最终固化为 checklist 的 5 个核心坑。5.1 坑UTF-8 源码文件读取时 BOM 导致首字符错位现象脚本第一行local a 1编译失败报错Unexpected token l at line 1, column 1但肉眼确认无空格。原因Windows 记事本保存的 UTF-8 文件默认带 BOMEF BB BFFile.ReadAllText(path)会把 BOM 当作普通字符读入导致source[0]是而非l。解决读取时强制指定编码File.ReadAllText(path, Encoding.UTF8).NET Core 3.0 默认已忽略 BOM但 .NET Framework 4.7.2 及以下仍需显式指定更稳妥做法用StreamReader并设置detectEncodingFromByteOrderMarks: true在 Lexer 开头加校验if (source.Length 0 source[0] \uFEFF) source source.Slice(1);。5.2 坑Pratt 解析器中^右结合性未正确实现现象2 ^ 3 ^ 4解析为(2 ^ 3) ^ 4 8 ^ 4 4096但 Lua 正确结果是2 ^ (3 ^ 4) 2 ^ 81 ≈ 2.4e24。原因Caret的rbp设为 60 后ParseExpression(60)会立即停止递归导致右操作数未被当作高优先级表达式解析。解决Caret的led函数中调用ParseExpression(60)而非rbp作为右操作数即right ParseExpression(60)或更规范做法为Caret单独设rbp 60并在ParseInfix中对Caret特殊处理强制right ParseExpression(61)高于自身。5.3 坑local x x 1未报错运行时NullReferenceException现象编译成功但执行时报Object reference not set to an instance of an object。原因语义分析只检查变量是否声明未区分“声明”和“初始化”。local x x 1中x在右侧是未初始化的局部变量应视为null但 Lua 规定此时x为nil而nil 1应报错。解决在LocalDeclarationNode的Expressions遍历时对每个IdentifierNode检查其是否在当前 scope 中已声明且已初始化更佳方案引入Definite Assignment Analysis跟踪每个局部变量的赋值状态未赋值就引用则报Variable x is used before being assigned。5.4 坑for i 1, 10 do ... end生成的FORPREP/FORLOOP指令跳转地址错乱现象循环体执行一次后直接跳出或无限循环。原因FORPREP指令需跳转到循环体末尾的FORLOOP但代码生成时FORLOOP指令尚未 emit无法预知其偏移。若用占位符 回填易因指令长度变化如Ldc_I4_SvsLdc_I4导致 offset 偏移。解决改用两趟生成第一趟收集所有跳转目标 label第二趟 emit 并计算真实 offset或用ILGenerator.MarkLabel()ILGenerator.Emit(OpCodes.Br_S, label)让 .NET 运行时自动 patch最简实践对for循环生成标准for (int i 1; i 10; i) { ... }的 IL绕过 Lua VM 指令模拟。5.5 坑Dynamic本文还有配套的精品资源点击获取
返回列表