SEARCH
NEW RPMS
DIRECTORIES
ABOUT
FAQ
VARIOUS
BLOG

BotDetect - Real-Time Bot Detection API
 
 

MAN page from Old RedHat 5.X flex-2.5.4a-3.i386.rpm

FLEX

Section: User Commands (1)
Updated: April 1995
Index 

NAME

flex - fast lexical analyzer generator 

SYNOPSIS

flex[-bcdfhilnpstvwBFILTV78+? -C[aefFmr] -ooutput -Pprefix -Sskeleton][--help --version][filename ...] 

OVERVIEW

This manual describesflex,a tool for generating programs that perform pattern-matching on text. Themanual includes both tutorial and reference sections:
    Description        a brief overview of the tool    Some Simple Examples    Format Of The Input File    Patterns        the extended regular expressions used by flex    How The Input Is Matched        the rules for determining what has been matched    Actions        how to specify what to do when a pattern is matched    The Generated Scanner        details regarding the scanner that flex produces;        how to control the input source    Start Conditions        introducing context into your scanners, and        managing "mini-scanners"    Multiple Input Buffers        how to manipulate multiple input sources; how to        scan from strings instead of files    End-of-file Rules        special rules for matching the end of the input    Miscellaneous Macros        a summary of macros available to the actions    Values Available To The User        a summary of values available to the actions    Interfacing With Yacc        connecting flex scanners together with yacc parsers    Options        flex command-line options, and the "%option"        directive    Performance Considerations        how to make your scanner go as fast as possible    Generating C++ Scanners        the (experimental) facility for generating C++        scanner classes    Incompatibilities With Lex And POSIX        how flex differs from AT&T lex and the POSIX lex        standard    Diagnostics        those error messages produced by flex (or scanners        it generates) whose meanings might not be apparent    Files        files used by flex    Deficiencies / Bugs        known problems with flex    See Also        other documentation, related tools    Author        includes contact information
 

DESCRIPTION

flexis a tool for generatingscanners:programs which recognized lexical patterns in text.flexreadsthe given input files, or its standard input if no file names are given,for a description of a scanner to generate. The description is inthe form of pairsof regular expressions and C code, calledrules. flexgenerates as output a C source file,lex.yy.c,which defines a routineyylex().This file is compiled and linked with the-lfllibrary to produce an executable. When the executable is run,it analyzes its input for occurrencesof the regular expressions. Whenever it finds one, it executesthe corresponding C code. 

SOME SIMPLE EXAMPLES

First some simple examples to get the flavor of how one usesflex.The followingflexinput specifies a scanner which whenever it encounters the string"username" will replace it with the user's login name:

    %%    username    printf( "%s", getlogin() );
By default, any text not matched by aflexscanneris copied to the output, so the net effect of this scanner isto copy its input file to its output with each occurrenceof "username" expanded.In this input, there is just one rule. "username" is thepatternand the "printf" is theaction.The "%%" marks the beginning of the rules.

Here's another simple example:

            int num_lines = 0, num_chars = 0;    %%    \n      ++num_lines; ++num_chars;    .       ++num_chars;    %%    main()            {            yylex();            printf( "# of lines = %d, # of chars = %d\n",                    num_lines, num_chars );            }
This scanner counts the number of characters and the numberof lines in its input (it produces no output other than thefinal report on the counts). The first linedeclares two globals, "num_lines" and "num_chars", which are accessibleboth insideyylex()and in themain()routine declared after the second "%%". There are two rules, onewhich matches a newline ("\n") and increments both the line count andthe character count, and one which matches any character other thana newline (indicated by the "." regular expression).

A somewhat more complicated example:

    /* scanner for a toy Pascal-like language */    %{    /* need this for the call to atof() below */    #include <math.h>    %}    DIGIT    [0-9]    ID       [a-z][a-z0-9]*    %%    {DIGIT}+    {                printf( "An integer: %s (%d)\n", yytext,                        atoi( yytext ) );                }    {DIGIT}+"."{DIGIT}*        {                printf( "A float: %s (%g)\n", yytext,                        atof( yytext ) );                }    if|then|begin|end|procedure|function        {                printf( "A keyword: %s\n", yytext );                }    {ID}        printf( "An identifier: %s\n", yytext );    "+"|"-"|"*"|"/"   printf( "An operator: %s\n", yytext );    "{"[^}\n]*"}"     /* eat up one-line comments */    [ \t\n]+          /* eat up whitespace */    .           printf( "Unrecognized character: %s\n", yytext );    %%    main( argc, argv )    int argc;    char **argv;        {        ++argv, --argc;  /* skip over program name */        if ( argc > 0 )                yyin = fopen( argv[0], "r" );        else                yyin = stdin;                yylex();        }
This is the beginnings of a simple scanner for a language likePascal. It identifies different types oftokensand reports on what it has seen.

The details of this example will be explained in the followingsections. 

FORMAT OF THE INPUT FILE

Theflexinput file consists of three sections, separated by a line with just%%in it:
    definitions    %%    rules    %%    user code
Thedefinitionssection contains declarations of simplenamedefinitions to simplify the scanner specification, and declarations ofstart conditions,which are explained in a later section.

Name definitions have the form:

    name definition
The "name" is a word beginning with a letter or an underscore ('_')followed by zero or more letters, digits, '_', or '-' (dash).The definition is taken to begin at the first non-white-space characterfollowing the name and continuing to the end of the line.The definition can subsequently be referred to using "{name}", whichwill expand to "(definition)". For example,
    DIGIT    [0-9]    ID       [a-z][a-z0-9]*
defines "DIGIT" to be a regular expression which matches asingle digit, and"ID" to be a regular expression which matches a letterfollowed by zero-or-more letters-or-digits.A subsequent reference to
    {DIGIT}+"."{DIGIT}*
is identical to
    ([0-9])+"."([0-9])*
and matches one-or-more digits followed by a '.' followedby zero-or-more digits.

Therulessection of theflexinput contains a series of rules of the form:

    pattern   action
where the pattern must be unindented and the action must beginon the same line.

See below for a further description of patterns and actions.

Finally, the user code section is simply copied tolex.yy.cverbatim.It is used for companion routines which call or are calledby the scanner. The presence of this section is optional;if it is missing, the second%%in the input file may be skipped, too.

In the definitions and rules sections, anyindentedtext or text enclosed in%{and%}is copied verbatim to the output (with the %{}'s removed).The %{}'s must appear unindented on lines by themselves.

In the rules section,any indented or %{} text appearing before thefirst rule may be used to declare variableswhich are local to the scanning routine and (after the declarations)code which is to be executed whenever the scanning routine is entered.Other indented or %{} text in the rule section is still copied to the output,but its meaning is not well-defined and it may well cause compile-timeerrors (this feature is present forPOSIXcompliance; see below for other such features).

In the definitions section (but not in the rules section),an unindented comment (i.e., a linebeginning with "/*") is also copied verbatim to the output upto the next "*/". 

PATTERNS

The patterns in the input are written using an extended set of regularexpressions. These are:
    x          match the character 'x'    .          any character (byte) except newline    [xyz]      a "character class"; in this case, the pattern                 matches either an 'x', a 'y', or a 'z'    [abj-oZ]   a "character class" with a range in it; matches                 an 'a', a 'b', any letter from 'j' through 'o',                 or a 'Z'    [^A-Z]     a "negated character class", i.e., any character                 but those in the class.  In this case, any                 character EXCEPT an uppercase letter.    [^A-Z\n]   any character EXCEPT an uppercase letter or                 a newline    r*         zero or more r's, where r is any regular expression    r+         one or more r's    r?         zero or one r's (that is, "an optional r")    r{2,5}     anywhere from two to five r's    r{2,}      two or more r's    r{4}       exactly 4 r's    {name}     the expansion of the "name" definition               (see above)    "[xyz]\"foo"               the literal string: [xyz]"foo    \X         if X is an 'a', 'b', 'f', 'n', 'r', 't', or 'v',                 then the ANSI-C interpretation of \x.                 Otherwise, a literal 'X' (used to escape                 operators such as '*')    \0         a NUL character (ASCII code 0)    \123       the character with octal value 123    \x2a       the character with hexadecimal value 2a    (r)        match an r; parentheses are used to override                 precedence (see below)    rs         the regular expression r followed by the                 regular expression s; called "concatenation"    r|s        either an r or an s    r/s        an r but only if it is followed by an s.  The                 text matched by s is included when determining                 whether this rule is the "longest match",                 but is then returned to the input before                 the action is executed.  So the action only                 sees the text matched by r.  This type                 of pattern is called trailing context".                 (There are some combinations of r/s that flex                 cannot match correctly; see notes in the                 Deficiencies / Bugs section below regarding                 "dangerous trailing context".)    ^r         an r, but only at the beginning of a line (i.e.,                 which just starting to scan, or right after a                 newline has been scanned).    r$         an r, but only at the end of a line (i.e., just                 before a newline).  Equivalent to "r/\n".               Note that flex's notion of "newline" is exactly               whatever the C compiler used to compile flex               interprets '\n' as; in particular, on some DOS               systems you must either filter out \r's in the               input yourself, or explicitly use r/\r\n for "r$".    <s>r       an r, but only in start condition s (see                 below for discussion of start conditions)    <s1,s2,s3>r               same, but in any of start conditions s1,                 s2, or s3    <*>r       an r in any start condition, even an exclusive one.    <<EOF>>    an end-of-file    <s1,s2><<EOF>>               an end-of-file when in start condition s1 or s2
Note that inside of a character class, all regular expression operatorslose their special meaning except escape ('\') and the character classoperators, '-', ']', and, at the beginning of the class, '^'.

The regular expressions listed above are grouped according toprecedence, from highest precedence at the top to lowest at the bottom.Those grouped together have equal precedence. For example,

    foo|bar*
is the same as
    (foo)|(ba(r*))
since the '*' operator has higher precedence than concatenation,and concatenation higher than alternation ('|'). This patterntherefore matcheseitherthe string "foo"orthe string "ba" followed by zero-or-more r's.To match "foo" or zero-or-more "bar"'s, use:
    foo|(bar)*
and to match zero-or-more "foo"'s-or-"bar"'s:
    (foo|bar)*

In addition to characters and ranges of characters, character classescan also contain character classexpressions.These are expressions enclosed inside[:and:]delimiters (which themselves must appear between the '[' and ']' of thecharacter class; other elements may occur inside the character class, too).The valid expressions are:

    [:alnum:] [:alpha:] [:blank:]    [:cntrl:] [:digit:] [:graph:]    [:lower:] [:print:] [:punct:]    [:space:] [:upper:] [:xdigit:]
These expressions all designate a set of characters equivalent tothe corresponding standard CisXXXfunction. For example,[:alnum:]designates those characters for whichisalnum()returns true - i.e., any alphabetic or numeric.Some systems don't provideisblank(),so flex defines[:blank:]as a blank or a tab.

For example, the following character classes are all equivalent:

    [[:alnum:]]    [[:alpha:][:digit:]    [[:alpha:]0-9]    [a-zA-Z0-9]
If your scanner is case-insensitive (the-iflag), then[:upper:]and[:lower:]are equivalent to[:alpha:].

Some notes on patterns:

-
A negated character class such as the example "[^A-Z]"abovewill match a newlineunless "\n" (or an equivalent escape sequence) is one of thecharacters explicitly present in the negated character class(e.g., "[^A-Z\n]"). This is unlike how many other regularexpression tools treat negated character classes, but unfortunatelythe inconsistency is historically entrenched.Matching newlines means that a pattern like [^"]* can match the entireinput unless there's another quote in the input.
-
A rule can have at most one instance of trailing context (the '/' operatoror the '$' operator). The start condition, '^', and "<<EOF>>" patternscan only occur at the beginning of a pattern, and, as well as with '/' and '$',cannot be grouped inside parentheses. A '^' which does not occur atthe beginning of a rule or a '$' which does not occur at the end ofa rule loses its special properties and is treated as a normal character.
The following are illegal:
    foo/bar$    <sc1>foo<sc2>bar
Note that the first of these, can be written "foo/bar\n".
The following will result in '$' or '^' being treated as a normal character:
    foo|(bar$)    foo|^bar
If what's wanted is a "foo" or a bar-followed-by-a-newline, the followingcould be used (the special '|' action is explained below):
    foo      |    bar$     /* action goes here */
A similar trick will work for matching a foo or abar-at-the-beginning-of-a-line.
 

HOW THE INPUT IS MATCHED

When the generated scanner is run, it analyzes its input lookingfor strings which match any of its patterns. If it finds more thanone match, it takes the one matching the most text (for trailingcontext rules, this includes the length of the trailing part, eventhough it will then be returned to the input). If it finds twoor more matches of the same length, therule listed first in theflexinput file is chosen.

Once the match is determined, the text corresponding to the match(called thetoken)is made available in the global character pointeryytext,and its length in the global integeryyleng.Theactioncorresponding to the matched pattern is then executed (a moredetailed description of actions follows), and then the remaininginput is scanned for another match.

If no match is found, then thedefault ruleis executed: the next character in the input is considered matched andcopied to the standard output. Thus, the simplest legalflexinput is:

    %%
which generates a scanner that simply copies its input (one characterat a time) to its output.

Note thatyytextcan be defined in two different ways: either as a characterpointeror as a characterarray.You can control which definitionflexuses by including one of the special directives%pointeror%arrayin the first (definitions) section of your flex input. The default is%pointer,unless you use the-llex compatibility option, in which caseyytextwill be an array.The advantage of using%pointeris substantially faster scanning and no buffer overflow when matchingvery large tokens (unless you run out of dynamic memory). The disadvantageis that you are restricted in how your actions can modifyyytext(see the next section), and calls to theunput()function destroys the present contents ofyytext,which can be a considerable porting headache when moving between differentlexversions.

The advantage of%arrayis that you can then modifyyytextto your heart's content, and calls tounput()do not destroyyytext(see below). Furthermore, existinglexprograms sometimes accessyytextexternally using declarations of the form:

    extern char yytext[];
This definition is erroneous when used with%pointer,but correct for%array.

%arraydefinesyytextto be an array ofYYLMAXcharacters, which defaults to a fairly large value. You can changethe size by simply #define'ingYYLMAXto a different value in the first section of yourflexinput. As mentioned above, with%pointeryytext grows dynamically to accommodate large tokens. While this means your%pointerscanner can accommodate very large tokens (such as matching entire blocksof comments), bear in mind that each time the scanner must resizeyytextit also must rescan the entire token from the beginning, so matching suchtokens can prove slow.yytextpresently doesnotdynamically grow if a call tounput()results in too much text being pushed back; instead, a run-time error results.

Also note that you cannot use%arraywith C++ scanner classes(thec++option; see below). 

ACTIONS

Each pattern in a rule has a corresponding action, which can be anyarbitrary C statement. The pattern ends at the first non-escapedwhitespace character; the remainder of the line is its action. If theaction is empty, then when the pattern is matched the input tokenis simply discarded. For example, here is the specification for a programwhich deletes all occurrences of "zap me" from its input:
    %%    "zap me"
(It will copy all other characters in the input to the output sincethey will be matched by the default rule.)

Here is a program which compresses multiple blanks and tabs down toa single blank, and throws away whitespace found at the end of a line:

    %%    [ \t]+        putchar( ' ' );    [ \t]+$       /* ignore this token */

If the action contains a '{', then the action spans till the balancing '}'is found, and the action may cross multiple lines.flex knows about C strings and comments and won't be fooled by braces foundwithin them, but also allows actions to begin with%{and will consider the action to be all the text up to the next%}(regardless of ordinary braces inside the action).

An action consisting solely of a vertical bar ('|') means "same asthe action for the next rule." See below for an illustration.

Actions can include arbitrary C code, includingreturnstatements to return a value to whatever routine calledyylex().Each timeyylex()is called it continues processing tokens from where it last leftoff until it either reachesthe end of the file or executes a return.

Actions are free to modifyyytextexcept for lengthening it (addingcharacters to its end--these will overwrite later characters in theinput stream). This however does not apply when using%array(see above); in that case,yytextmay be freely modified in any way.

Actions are free to modifyyylengexcept they should not do so if the action also includes use ofyymore()(see below).

There are a number of special directives which can be included withinan action:

-
ECHOcopies yytext to the scanner's output.
-
BEGINfollowed by the name of a start condition places the scanner in thecorresponding start condition (see below).
-
REJECTdirects the scanner to proceed on to the "second best" rule which matched theinput (or a prefix of the input). The rule is chosen as describedabove in "How the Input is Matched", andyytextandyylengset up appropriately.It may either be one which matched as much textas the originally chosen rule but came later in theflexinput file, or one which matched less text.For example, the following will both count thewords in the input and call the routine special() whenever "frob" is seen:
            int word_count = 0;    %%    frob        special(); REJECT;    [^ \t\n]+   ++word_count;
Without theREJECT,any "frob"'s in the input would not be counted as words, since thescanner normally executes only one action per token.MultipleREJECT'sare allowed, each one finding the next best choice to the currentlyactive rule. For example, when the following scanner scans the token"abcd", it will write "abcdabcaba" to the output:
    %%    a        |    ab       |    abc      |    abcd     ECHO; REJECT;    .|\n     /* eat up any unmatched character */
(The first three rules share the fourth's action since they usethe special '|' action.)REJECTis a particularly expensive feature in terms of scanner performance;if it is used inanyof the scanner's actions it will slow downallof the scanner's matching. Furthermore,REJECTcannot be used with the-Cfor-CFoptions (see below).
Note also that unlike the other special actions,REJECTis abranch;code immediately following it in the action willnotbe executed.
-
yymore()tells the scanner that the next time it matches a rule, the correspondingtoken should beappendedonto the current value ofyytextrather than replacing it. For example, given the input "mega-kludge"the following will write "mega-mega-kludge" to the output:
    %%    mega-    ECHO; yymore();    kludge   ECHO;
First "mega-" is matched and echoed to the output. Then "kludge"is matched, but the previous "mega-" is still hanging around at thebeginning ofyytextso theECHOfor the "kludge" rule will actually write "mega-kludge".

Two notes regarding use ofyymore().First,yymore()depends on the value ofyylengcorrectly reflecting the size of the current token, so you must notmodifyyylengif you are usingyymore().Second, the presence ofyymore()in the scanner's action entails a minor performance penalty in thescanner's matching speed.

-
yyless(n)returns all but the firstncharacters of the current token back to the input stream, where theywill be rescanned when the scanner looks for the next match.yytextandyylengare adjusted appropriately (e.g.,yylengwill now be equal ton). For example, on the input "foobar" the following will write out"foobarbar":
    %%    foobar    ECHO; yyless(3);    [a-z]+    ECHO;
An argument of 0 toyylesswill cause the entire current input string to be scanned again. Unless you'vechanged how the scanner will subsequently process its input (usingBEGIN,for example), this will result in an endless loop.

Note thatyylessis a macro and can only be used in the flex input file, not fromother source files.

-
unput(c)puts the charactercback onto the input stream. It will be the next character scanned.The following action will take the current token and cause itto be rescanned enclosed in parentheses.
    {    int i;    /* Copy yytext because unput() trashes yytext */    char *yycopy = strdup( yytext );    unput( ')' );    for ( i = yyleng - 1; i >= 0; --i )        unput( yycopy[i] );    unput( '(' );    free( yycopy );    }
Note that since eachunput()puts the given character back at thebeginningof the input stream, pushing back strings must be done back-to-front.

An important potential problem when usingunput()is that if you are using%pointer(the default), a call tounput()destroysthe contents ofyytext,starting with its rightmost character and devouring one character tothe left with each call. If you need the value of yytext preservedafter a call tounput()(as in the above example),you must either first copy it elsewhere, or build your scanner using%arrayinstead (see How The Input Is Matched).

Finally, note that you cannot put backEOFto attempt to mark the input stream with an end-of-file.

-
input()reads the next character from the input stream. For example,the following is one way to eat up C comments:
    %%    "/*"        {                register int c;                for ( ; ; )                    {                    while ( (c = input()) != '*' &&                            c != EOF )                        ;    /* eat up text of comment */                    if ( c == '*' )                        {                        while ( (c = input()) == '*' )                            ;                        if ( c == '/' )                            break;    /* found the end */                        }                    if ( c == EOF )                        {                        error( "EOF in comment" );                        break;                        }                    }                }
(Note that if the scanner is compiled usingC++,theninput()is instead referred to asyyinput(),in order to avoid a name clash with theC++stream by the name ofinput.)
-
YY_FLUSH_BUFFERflushes the scanner's internal bufferso that the next time the scanner attempts to match a token, it willfirst refill the buffer usingYY_INPUT(see The Generated Scanner, below). This action is a special caseof the more generalyy_flush_buffer()function, described below in the section Multiple Input Buffers.
-
yyterminate()can be used in lieu of a return statement in an action. It terminatesthe scanner and returns a 0 to the scanner's caller, indicating "all done".By default,yyterminate()is also called when an end-of-file is encountered. It is a macro andmay be redefined.
 

THE GENERATED SCANNER

The output offlexis the filelex.yy.c,which contains the scanning routineyylex(),a number of tables used by it for matching tokens, and a numberof auxiliary routines and macros. By default,yylex()is declared as follows:
    int yylex()        {        ... various definitions and the actions in here ...        }
(If your environment supports function prototypes, then it willbe "int yylex( void )".) This definition may be changed by definingthe "YY_DECL" macro. For example, you could use:
    #define YY_DECL float lexscan( a, b ) float a, b;
to give the scanning routine the namelexscan,returning a float, and taking two floats as arguments. Note thatif you give arguments to the scanning routine using aK&R-style/non-prototyped function declaration, you must terminatethe definition with a semi-colon (;).

Wheneveryylex()is called, it scans tokens from the global input fileyyin(which defaults to stdin). It continues until it either reachesan end-of-file (at which point it returns the value 0) orone of its actions executes areturnstatement.

If the scanner reaches an end-of-file, subsequent calls are undefinedunless eitheryyinis pointed at a new input file (in which case scanning continues fromthat file), oryyrestart()is called.yyrestart()takes one argument, aFILE *pointer (which can be nil, if you've set upYY_INPUTto scan from a source other thanyyin),and initializesyyinfor scanning from that file. Essentially there is no difference betweenjust assigningyyinto a new input file or usingyyrestart()to do so; the latter is available for compatibility with previous versionsofflex,and because it can be used to switch input files in the middle of scanning.It can also be used to throw away the current input buffer, by callingit with an argument ofyyin;but better is to useYY_FLUSH_BUFFER(see above).Note thatyyrestart()doesnotreset the start condition toINITIAL(see Start Conditions, below).

Ifyylex()stops scanning due to executing areturnstatement in one of the actions, the scanner may then be called again and itwill resume scanning where it left off.

By default (and for purposes of efficiency), the scanner usesblock-reads rather than simplegetc()calls to read characters fromyyin.The nature of how it gets its input can be controlled by defining theYY_INPUTmacro.YY_INPUT's calling sequence is "YY_INPUT(buf,result,max_size)". Itsaction is to place up tomax_sizecharacters in the character arraybufand return in the integer variableresulteither thenumber of characters read or the constant YY_NULL (0 on Unix systems)to indicate EOF. The default YY_INPUT reads from theglobal file-pointer "yyin".

A sample definition of YY_INPUT (in the definitionssection of the input file):

    %{    #define YY_INPUT(buf,result,max_size) \        { \        int c = getchar(); \        result = (c == EOF) ? YY_NULL : (buf[0] = c, 1); \        }    %}
This definition will change the input processing to occurone character at a time.

When the scanner receives an end-of-file indication from YY_INPUT,it then checks theyywrap()function. Ifyywrap()returns false (zero), then it is assumed that thefunction has gone ahead and set upyyinto point to another input file, and scanning continues. If it returnstrue (non-zero), then the scanner terminates, returning 0 to itscaller. Note that in either case, the start condition remains unchanged;it doesnotrevert toINITIAL.

If you do not supply your own version ofyywrap(),then you must either use%option noyywrap(in which case the scanner behaves as thoughyywrap()returned 1), or you must link with-lflto obtain the default version of the routine, which always returns 1.

Three routines are available for scanning from in-memory buffers ratherthan files:yy_scan_string(), yy_scan_bytes(),andyy_scan_buffer().See the discussion of them below in the section Multiple Input Buffers.

The scanner writes itsECHOoutput to theyyoutglobal (default, stdout), which may be redefined by the user simplyby assigning it to some otherFILEpointer. 

START CONDITIONS

flexprovides a mechanism for conditionally activating rules. Any rulewhose pattern is prefixed with "<sc>" will only be active whenthe scanner is in the start condition named "sc". For example,
    <STRING>[^"]*        { /* eat up the string body ... */                ...                }
will be active only when the scanner is in the "STRING" startcondition, and
    <INITIAL,STRING,QUOTE>\.        { /* handle an escape ... */                ...                }
will be active only when the current start condition iseither "INITIAL", "STRING", or "QUOTE".

Start conditionsare declared in the definitions (first) section of the inputusing unindented lines beginning with either%sor%xfollowed by a list of names.The former declaresinclusivestart conditions, the latterexclusivestart conditions. A start condition is activated using theBEGINaction. Until the nextBEGINaction is executed, rules with the given startcondition will be active andrules with other start conditions will be inactive.If the start condition isinclusive,then rules with no start conditions at all will also be active.If it isexclusive,thenonlyrules qualified with the start condition will be active.A set of rules contingent on the same exclusive start conditiondescribe a scanner which is independent of any of the other rules in theflexinput. Because of this,exclusive start conditions make it easy to specify "mini-scanners"which scan portions of the input that are syntactically differentfrom the rest (e.g., comments).

If the distinction between inclusive and exclusive start conditionsis still a little vague, here's a simple example illustrating theconnection between the two. The set of rules:

    %s example    %%    <example>foo   do_something();    bar            something_else();
is equivalent to
    %x example    %%    <example>foo   do_something();    <INITIAL,example>bar    something_else();
Without the<INITIAL,example>qualifier, thebarpattern in the second example wouldn't be active (i.e., couldn't match)when in start conditionexample.If we just used<example>to qualifybar,though, then it would only be active inexampleand not inINITIAL,while in the first example it's active in both, because in the firstexample theexamplestartion condition is aninclusive(%s)start condition.

Also note that the special start-condition specifier<*>matches every start condition. Thus, the above example could alsohave been written;

    %x example    %%    <example>foo   do_something();    <*>bar    something_else();

The default rule (toECHOany unmatched character) remains active in start conditions. Itis equivalent to:

    <*>.|\n     ECHO;

BEGIN(0)returns to the original state where only the rules withno start conditions are active. This state can also bereferred to as the start-condition "INITIAL", soBEGIN(INITIAL)is equivalent toBEGIN(0).(The parentheses around the start condition name are not required butare considered good style.)

BEGINactions can also be given as indented code at the beginningof the rules section. For example, the following will causethe scanner to enter the "SPECIAL" start condition wheneveryylex()is called and the global variableenter_specialis true:

            int enter_special;    %x SPECIAL    %%            if ( enter_special )                BEGIN(SPECIAL);    <SPECIAL>blahblahblah    ...more rules follow...

To illustrate the uses of start conditions,here is a scanner which provides two different interpretationsof a string like "123.456". By default it will treat it asthree tokens, the integer "123", a dot ('.'), and the integer "456".But if the string is preceded earlier in the line by the string"expect-floats"it will treat it as a single token, the floating-point number123.456:

    %{    #include <math.h>    %}    %s expect    %%    expect-floats        BEGIN(expect);    <expect>[0-9]+"."[0-9]+      {                printf( "found a float, = %f\n",                        atof( yytext ) );                }    <expect>\n           {                /* that's the end of the line, so                 * we need another "expect-number"                 * before we'll recognize any more                 * numbers                 */                BEGIN(INITIAL);                }    [0-9]+      {                printf( "found an integer, = %d\n",                        atoi( yytext ) );                }    "."         printf( "found a dot\n" );
Here is a scanner which recognizes (and discards) C comments whilemaintaining a count of the current input line.
    %x comment    %%            int line_num = 1;    "/*"         BEGIN(comment);    <comment>[^*\n]*        /* eat anything that's not a '*' */    <comment>"*"+[^*/\n]*   /* eat up '*'s not followed by '/'s */    <comment>\n             ++line_num;    <comment>"*"+"/"        BEGIN(INITIAL);
This scanner goes to a bit of trouble to match as muchtext as possible with each rule. In general, when attempting to writea high-speed scanner try to match as much possible in each rule, asit's a big win.

Note that start-conditions names are really integer values andcan be stored as such. Thus, the above could be extended in thefollowing fashion:

    %x comment foo    %%            int line_num = 1;            int comment_caller;    "/*"         {                 comment_caller = INITIAL;                 BEGIN(comment);                 }    ...    <foo>"/*"    {                 comment_caller = foo;                 BEGIN(comment);                 }    <comment>[^*\n]*        /* eat anything that's not a '*' */    <comment>"*"+[^*/\n]*   /* eat up '*'s not followed by '/'s */    <comment>\n             ++line_num;    <comment>"*"+"/"        BEGIN(comment_caller);
Furthermore, you can access the current start condition usingthe integer-valuedYY_STARTmacro. For example, the above assignments tocomment_callercould instead be written
    comment_caller = YY_START;
Flex providesYYSTATEas an alias forYY_START(since that is what's used by AT&Tlex).

Note that start conditions do not have their own name-space; %s's and %x'sdeclare names in the same fashion as #define's.

Finally, here's an example of how to match C-style quoted strings usingexclusive start conditions, including expanded escape sequences (butnot including checking for a string that's too long):

    %x str    %%            char string_buf[MAX_STR_CONST];            char *string_buf_ptr;    \"      string_buf_ptr = string_buf; BEGIN(str);    <str>\"        { /* saw closing quote - all done */            BEGIN(INITIAL);            *string_buf_ptr = '\0';            /* return string constant token type and             * value to parser             */            }    <str>\n        {            /* error - unterminated string constant */            /* generate error message */            }    <str>\\[0-7]{1,3} {            /* octal escape sequence */            int result;            (void) sscanf( yytext + 1, "%o", &result );            if ( result > 0xff )                    /* error, constant is out-of-bounds */            *string_buf_ptr++ = result;            }    <str>\\[0-9]+ {            /* generate error - bad escape sequence; something             * like '\48' or '\0777777'             */            }    <str>\\n  *string_buf_ptr++ = '\n';    <str>\\t  *string_buf_ptr++ = '\t';    <str>\\r  *string_buf_ptr++ = '\r';    <str>\\b  *string_buf_ptr++ = '\b';    <str>\\f  *string_buf_ptr++ = '\f';    <str>\\(.|\n)  *string_buf_ptr++ = yytext[1];    <str>[^\\\n\"]+        {            char *yptr = yytext;            while ( *yptr )                    *string_buf_ptr++ = *yptr++;            }

Often, such as in some of the examples above, you wind up writing awhole bunch of rules all preceded by the same start condition(s). Flexmakes this a little easier and cleaner by introducing a notion ofstart conditionscope.A start condition scope is begun with:

    <SCs>{
whereSCsis a list of one or more start conditions. Inside the start conditionscope, every rule automatically has the prefix<SCs>applied to it, until a'}'which matches the initial'{'.So, for example,
    <ESC>{        "\\n"   return '\n';        "\\r"   return '\r';        "\\f"   return '\f';        "\\0"   return '\0';    }
is equivalent to:
    <ESC>"\\n"  return '\n';    <ESC>"\\r"  return '\r';    <ESC>"\\f"  return '\f';    <ESC>"\\0"  return '\0';
Start condition scopes may be nested.

Three routines are available for manipulating stacks of start conditions:

void yy_push_state(int new_state)
pushes the current start condition onto the top of the start conditionstack and switches tonew_stateas though you had usedBEGIN new_state(recall that start condition names are also integers).
void yy_pop_state()
pops the top of the stack and switches to it viaBEGIN.
int yy_top_state()
returns the top of the stack without altering the stack's contents.

The start condition stack grows dynamically and so has no built-insize limitation. If memory is exhausted, program execution aborts.

To use start condition stacks, your scanner must include a%option stackdirective (see Options below). 

MULTIPLE INPUT BUFFERS

Some scanners (such as those which support "include" files)require reading from several input streams. Asflexscanners do a large amount of buffering, one cannot controlwhere the next input will be read from by simply writing aYY_INPUTwhich is sensitive to the scanning context.YY_INPUTis only called when the scanner reaches the end of its buffer, whichmay be a long time after scanning a statement such as an "include"which requires switching the input source.

To negotiate these sorts of problems,flexprovides a mechanism for creating and switching between multipleinput buffers. An input buffer is created by using:

    YY_BUFFER_STATE yy_create_buffer( FILE *file, int size )
which takes aFILEpointer and a size and creates a buffer associated with the givenfile and large enough to holdsizecharacters (when in doubt, useYY_BUF_SIZEfor the size). It returns aYY_BUFFER_STATEhandle, which may then be passed to other routines (see below). TheYY_BUFFER_STATEtype is a pointer to an opaquestruct yy_buffer_statestructure, so you may safely initialize YY_BUFFER_STATE variables to((YY_BUFFER_STATE) 0)if you wish, and also refer to the opaque structure in order tocorrectly declare input buffers in source files other than thatof your scanner. Note that theFILEpointer in the call toyy_create_bufferis only used as the value ofyyinseen byYY_INPUT;if you redefineYY_INPUTso it no longer usesyyin,then you can safely pass a nilFILEpointer toyy_create_buffer.You select a particular buffer to scan from using:
    void yy_switch_to_buffer( YY_BUFFER_STATE new_buffer )
switches the scanner's input buffer so subsequent tokens willcome fromnew_buffer.Note thatyy_switch_to_buffer()may be used by yywrap() to set things up for continued scanning, insteadof opening a new file and pointingyyinat it. Note also that switching input sources via eitheryy_switch_to_buffer()oryywrap()doesnotchange the start condition.
    void yy_delete_buffer( YY_BUFFER_STATE buffer )
is used to reclaim the storage associated with a buffer. (buffercan be nil, in which case the routine does nothing.)You can also clear the current contents of a buffer using:
    void yy_flush_buffer( YY_BUFFER_STATE buffer )
This function discards the buffer's contents,so the next time the scanner attempts to match a token from thebuffer, it will first fill the buffer anew usingYY_INPUT.

yy_new_buffer()is an alias foryy_create_buffer(),provided for compatibility with the C++ use ofnewanddeletefor creating and destroying dynamic objects.

Finally, theYY_CURRENT_BUFFERmacro returns aYY_BUFFER_STATEhandle to the current buffer.

Here is an example of using these features for writing a scannerwhich expands include files (the<<EOF>>feature is discussed below):

    /* the "incl" state is used for picking up the name     * of an include file     */    %x incl    %{    #define MAX_INCLUDE_DEPTH 10    YY_BUFFER_STATE include_stack[MAX_INCLUDE_DEPTH];    int include_stack_ptr = 0;    %}    %%    include             BEGIN(incl);    [a-z]+              ECHO;    [^a-z\n]*\n?        ECHO;    <incl>[ \t]*      /* eat the whitespace */    <incl>[^ \t\n]+   { /* got the include file name */            if ( include_stack_ptr >= MAX_INCLUDE_DEPTH )                {                fprintf( stderr, "Includes nested too deeply" );                exit( 1 );                }            include_stack[include_stack_ptr++] =                YY_CURRENT_BUFFER;            yyin = fopen( yytext, "r" );            if ( ! yyin )                error( ... );            yy_switch_to_buffer(                yy_create_buffer( yyin, YY_BUF_SIZE ) );            BEGIN(INITIAL);            }    <<EOF>> {            if ( --include_stack_ptr < 0 )                {                yyterminate();                }            else                {                yy_delete_buffer( YY_CURRENT_BUFFER );                yy_switch_to_buffer(                     include_stack[include_stack_ptr] );                }            }
Three routines are available for setting up input buffers forscanning in-memory strings instead of files. All of them createa new input buffer for scanning the string, and return a correspondingYY_BUFFER_STATEhandle (which you should delete withyy_delete_buffer()when done with it). They also switch to the new buffer usingyy_switch_to_buffer(),so the next call toyylex()will start scanning the string.
yy_scan_string(const char *str)
scans a NUL-terminated string.
yy_scan_bytes(const char *bytes, int len)
scanslenbytes (including possibly NUL's)starting at locationbytes.

Note that both of these functions create and scan acopyof the string or bytes. (This may be desirable, sinceyylex()modifies the contents of the buffer it is scanning.) You can avoid thecopy by using:

yy_scan_buffer(char *base, yy_size_t size)
which scans in place the buffer starting atbase,consisting ofsizebytes, the last two bytes of whichmustbeYY_END_OF_BUFFER_CHAR(ASCII NUL).These last two bytes are not scanned; thus, scanningconsists ofbase[0]throughbase[size-2],inclusive.
If you fail to set upbasein this manner (i.e., forget the final twoYY_END_OF_BUFFER_CHARbytes), thenyy_scan_buffer()returns a nil pointer instead of creating a new input buffer.
The typeyy_size_tis an integral type to which you can cast an integer expressionreflecting the size of the buffer.
 

END-OF-FILE RULES

The special rule "<<EOF>>" indicatesactions which are to be taken when an end-of-file isencountered and yywrap() returns non-zero (i.e., indicatesno further files to process). The action must finishby doing one of four things:
-
assigningyyinto a new input file (in previous versions of flex, after doing theassignment you had to call the special actionYY_NEW_FILE;this is no longer necessary);
-
executing areturnstatement;
-
executing the specialyyterminate()action;
-
or, switching to a new buffer usingyy_switch_to_buffer()as shown in the example above.

<<EOF>> rules may not be used with otherpatterns; they may only be qualified with a list of startconditions. If an unqualified <<EOF>> rule is given, itapplies toallstart conditions which do not already have <<EOF>> actions. Tospecify an <<EOF>> rule for only the initial start condition, use

    <INITIAL><<EOF>>

These rules are useful for catching things like unclosed comments.An example:

    %x quote    %%    ...other rules for dealing with quotes...    <quote><<EOF>>   {             error( "unterminated quote" );             yyterminate();             }    <<EOF>>  {             if ( *++filelist )                 yyin = fopen( *filelist, "r" );             else                yyterminate();             }
 

MISCELLANEOUS MACROS

The macroYY_USER_ACTIONcan be defined to provide an actionwhich is always executed prior to the matched rule's action. For example,it could be #define'd to call a routine to convert yytext to lower-case.WhenYY_USER_ACTIONis invoked, the variableyy_actgives the number of the matched rule (rules are numbered starting with 1).Suppose you want to profile how often each of your rules is matched. Thefollowing would do the trick:
    #define YY_USER_ACTION ++ctr[yy_act]
wherectris an array to hold the counts for the different rules. Note thatthe macroYY_NUM_RULESgives the total number of rules (including the default rule, even ifyou use-s),so a correct declaration forctris:
    int ctr[YY_NUM_RULES];

The macroYY_USER_INITmay be defined to provide an action which is always executed beforethe first scan (and before the scanner's internal initializations are done).For example, it could be used to call a routine to readin a data table or open a logging file.

The macroyy_set_interactive(is_interactive)can be used to control whether the current buffer is consideredinteractive.An interactive buffer is processed more slowly,but must be used when the scanner's input source is indeedinteractive to avoid problems due to waiting to fill buffers(see the discussion of the-Iflag below). A non-zero valuein the macro invocation marks the buffer as interactive, a zero value as non-interactive. Note that use of this macro overrides%option always-interactiveor%option never-interactive(see Options below).yy_set_interactive()must be invoked prior to beginning to scan the buffer that is(or is not) to be considered interactive.

The macroyy_set_bol(at_bol)can be used to control whether the current buffer's scanningcontext for the next token match is done as though at thebeginning of a line. A non-zero macro argument makes rules anchored with

The macroYY_AT_BOL()returns true if the next token scanned from the current bufferwill have '^' rules active, false otherwise.

In the generated scanner, the actions are all gathered in one largeswitch statement and separated usingYY_BREAK,which may be redefined. By default, it is simply a "break", to separateeach rule's action from the following rule's.RedefiningYY_BREAKallows, for example, C++ users to#define YY_BREAK to do nothing (while being very careful that everyrule ends with a "break" or a "return"!) to avoid suffering fromunreachable statement warnings where because a rule's action ends with"return", theYY_BREAKis inaccessible. 

VALUES AVAILABLE TO THE USER

This section summarizes the various values available to the userin the rule actions.
-
char *yytextholds the text of the current token. It may be modified but not lengthened(you cannot append characters to the end).
If the special directive%arrayappears in the first section of the scanner description, thenyytextis instead declaredchar yytext[YYLMAX],whereYYLMAXis a macro definition that you can redefine in the first sectionif you don't like the default value (generally 8KB). Using%arrayresults in somewhat slower scanners, but the value ofyytextbecomes immune to calls toinput()andunput(),which potentially destroy its value whenyytextis a character pointer. The opposite of%arrayis%pointer,which is the default.
You cannot use%arraywhen generating C++ scanner classes(the-+flag).
-
int yylengholds the length of the current token.
-
FILE *yyinis the file which by defaultflexreads from. It may be redefined but doing so only makes sense beforescanning begins or after an EOF has been encountered. Changing it inthe midst of scanning will have unexpected results sinceflexbuffers its input; useyyrestart()instead.Once scanning terminates because an end-of-filehas been seen, you can assignyyinat the new input file and then call the scanner again to continue scanning.
-
void yyrestart( FILE *new_file )may be called to pointyyinat the new input file. The switch-over to the new file is immediate(any previously buffered-up input is lost). Note that callingyyrestart()withyyinas an argument thus throws away the current input buffer and continuesscanning the same input file.
-
FILE *yyoutis the file to whichECHOactions are done. It can be reassigned by the user.
-
YY_CURRENT_BUFFERreturns aYY_BUFFER_STATEhandle to the current buffer.
-
YY_STARTreturns an integer value corresponding to the current startcondition. You can subsequently use this value withBEGINto return to that start condition.
 

INTERFACING WITH YACC

One of the main uses offlexis as a companion to theyaccparser-generator.yaccparsers expect to call a routine namedyylex()to find the next input token. The routine is supposed toreturn the type of the next token as well as putting any associatedvalue in the globalyylval.To useflexwithyacc,one specifies the-doption toyaccto instruct it to generate the filey.tab.hcontaining definitions of all the%tokensappearing in theyaccinput. This file is then included in theflexscanner. For example, if one of the tokens is "TOK_NUMBER",part of the scanner might look like:
    %{    #include "y.tab.h"    %}    %%    [0-9]+        yylval = atoi( yytext ); return TOK_NUMBER;
 

OPTIONS

flexhas the following options:
-b
Generate backing-up information tolex.backup.This is a list of scanner states which require backing upand the input characters on which they do so. By adding rules onecan remove backing-up states. Ifallbacking-up states are eliminated and-Cfor-CFis used, the generated scanner will run faster (see the-pflag). Only users who wish to squeeze every last cycle out of theirscanners need worry about this option. (See the section on PerformanceConsiderations below.)
-c
is a do-nothing, deprecated option included for POSIX compliance.
-d
makes the generated scanner run indebugmode. Whenever a pattern is recognized and the globalyy_flex_debugis non-zero (which is the default),the scanner will write tostderra line of the form:
    --accepting rule at line 53 ("the matched text")
The line number refers to the location of the rule in the filedefining the scanner (i.e., the file that was fed to flex). Messagesare also generated when the scanner backs up, accepts thedefault rule, reaches the end of its input buffer (or encountersa NUL; at this point, the two look the same as far as the scanner's concerned),or reaches an end-of-file.
-f
specifiesfast scanner.No table compression is done and stdio is bypassed.The result is large but fast. This option is equivalent to-Cfr(see below).
-h
generates a "help" summary offlex'soptions tostdout and then exits.-?and--helpare synonyms for-h.
-i
instructsflexto generate acase-insensitivescanner. The case of letters given in theflexinput patterns willbe ignored, and tokens in the input will be matched regardless of case. Thematched text given inyytextwill have the preserved case (i.e., it will not be folded).
-l
turns on maximum compatibility with the original AT&Tleximplementation. Note that this does not meanfullcompatibility. Use of this option costs a considerable amount ofperformance, and it cannot be used with the-+, -f, -F, -Cf,or-CFoptions. For details on the compatibilities it provides, see the section"Incompatibilities With Lex And POSIX" below. This option also resultsin the nameYY_FLEX_LEX_COMPATbeing #define'd in the generated scanner.
-n
is another do-nothing, deprecated option included only forPOSIX compliance.
-p
generates a performance report to stderr. The reportconsists of comments regarding features of theflexinput file which will cause a serious loss of performance in the resultingscanner. If you give the flag twice, you will also get comments regardingfeatures that lead to minor performance losses.
Note that the use ofREJECT,%option yylineno,and variable trailing context (see the Deficiencies / Bugs section below)entails a substantial performance penalty; use ofyymore(),the^operator,and the-Iflag entail minor performance penalties.
-s
causes thedefault rule(that unmatched scanner input is echoed tostdout)to be suppressed. If the scanner encounters input that does notmatch any of its rules, it aborts with an error. This option isuseful for finding holes in a scanner's rule set.
-t
instructsflexto write the scanner it generates to standard output insteadoflex.yy.c.
-v
specifies thatflexshould write tostderra summary of statistics regarding the scanner it generates.Most of the statistics are meaningless to the casualflexuser, but the first line identifies the version offlex(same as reported by-V),and the next line the flags used when generating the scanner, includingthose that are on by default.
-w
suppresses warning messages.
-B
instructsflexto generate abatchscanner, the opposite ofinteractivescanners generated by-I(see below). In general, you use-Bwhen you arecertainthat your scanner will never be used interactively, and you want tosqueeze alittlemore performance out of it. If your goal is instead to squeeze out alotmore performance, you should be using the-Cfor-CFoptions (discussed below), which turn on-Bautomatically anyway.
-F
specifies that thefastscanner table representation should be used (and stdiobypassed). This representation isabout as fast as the full table representation(-f),and for some sets of patterns will be considerably smaller (and forothers, larger). In general, if the pattern set contains both "keywords"and a catch-all, "identifier" rule, such as in the set:
    "case"    return TOK_CASE;    "switch"  return TOK_SWITCH;    ...    "default" return TOK_DEFAULT;    [a-z]+    return TOK_ID;
then you're better off using the full table representation. If onlythe "identifier" rule is present and you then use a hash table or some suchto detect the keywords, you're better off using-F.
This option is equivalent to-CFr(see below). It cannot be used with-+.
-I
instructsflexto generate aninteractivescanner. An interactive scanner is one that only looks ahead to decidewhat token has been matched if it absolutely must. It turns out thatalways looking one extra character ahead, even if the scanner has alreadyseen enough text to disambiguate the current token, is a bit faster thanonly looking ahead when necessary. But scanners that always look aheadgive dreadful interactive performance; for example, when a user typesa newline, it is not recognized as a newline token until they enteranothertoken, which often means typing
 
ICM Bot detect detector