Reverse Engineering Course. Chapter 2

Reverse Engineering Tools. Static Analysis Tools
RE_TUTORIALRETOOLSCLIELFSTATIC ANALYSIS

Now that we can extract all the details about the file in a general way, it’s time to go deeper and look into the specifics for binaries. Some of the tools we’ve seen provide some details, but in this chapter, we’re going to go deep inside the binaries.

When it comes to analyzing a binary, researchers use two main groups of techniques known as static and dynamic analysis. At the end, as it happens with almost everything in this world, researchers end up using a hybrid approach, mixing both techniques and switching from one to the other as it better suits them. Both types of analysis are intended to extract different kinds of information, and, as you can imagine, there are countermeasure techniques used to make each one of those analyses much harder. So, what are those types of analysis about?

Well, in essence, it’s pretty easy. Anything you can do without running the specimen (let’s call the binary this way so it sounds more scientific) is static analysis. Everything you do running the specimen is dynamic analysis.

Let’s see some examples:

  • Disassembling the binary is static analysis
  • Decompiling the binary is static analysis
  • Debugging the binary is dynamic analysis
  • Monitoring the binary in a sandbox is dynamic analysis

This is just to give you a glimpse of what we’re talking about, but don’t worry, we’ll dive into the specific tools to do these things right now.

Static Analysis

Let’s start taking a look at the tools we can use to carry out our static analysis. There aren’t that many, but you really have to master them. Let’s start.

readelf your best friend

The tool readelf is an official member of any toolchain and it’s the ultimate reference about what the ELF file contains. It’s able to dump all ELF internal data structures, and each one of those structures provides us with very useful pieces of information. This tool is part of your Static Analysis as you do not have to run the binary to extract this information.

Let’s see the main information we can retrieve using readelf. Let’s start with the flag -a that stands for ALL. This is all I’ll say about it; this flag just dumps all available information, so this is usually one of the first things you do when starting the analysis of a binary. Instead of looking at the output of this flag, let’s check the output of the individual ones, which at the end will be part of the -a output.

ELF Header

The very first flag to use is -h that prints the information on the ELF header. This is the output when run against our Challenge01.

$ readelf -h challenge01.x86_64.bin
ELF Header:
  Magic:   7f 45 4c 46 02 01 01 00 00 00 00 00 00 00 00 00
  Class:                             ELF64
  Data:                              2's complement, little endian
  Version:                           1 (current)
  OS/ABI:                            UNIX - System V
  ABI Version:                       0
  Type:                              DYN (Position-Independent Executable file)
  Machine:                           Advanced Micro Devices X86-64
  Version:                           0x1
  Entry point address:               0x1090
  Start of program headers:          64 (bytes into file)
  Start of section headers:          14264 (bytes into file)
  Flags:                             0x0
  Size of this header:               64 (bytes)
  Size of program headers:           56 (bytes)
  Number of program headers:         14
  Size of section headers:           64 (bytes)
  Number of section headers:         31
  Section header string table index: 30

Some of this information we already know from tools like file, but readelf gave us much more details. In addition to all the information needed to extract all other information in the file, i.e., number of sections and its size, the offset where the section table is stored, the program header offset, and each entry size, etc., it also tells us a few other things:

  Class:                             ELF64
  Data:                              2's complement, little endian
  Type:                              DYN (Position-Independent Executable file)
  Entry point address:               0x1090

These elements tell us, in the order they appear above, if the binary is 32 or 64 bits, if it’s Little or Big Endian, if it’s static or dynamic, and the entry point of the program. The address in memory (or the offset from the base address when ASLR is enabled) of the first instruction that will be executed.

The First Executed Instruction

The first instruction executed by a program is a matter of a point of view. For a regular C program, the first instruction executed is the very first instruction in the main function. That’s reality for most regular programmers. However, that’s not true. For a system programmer—somebody dealing with compilers, linkers, binary manipulation tools, or malware analysis—the first instruction executed is the instruction at the _start label, which corresponds with the Entry point reported by readelf.

That’s the end of the story for most people; however, that’s not always the first instruction executed by a program. When the program is dynamic, the process is different. The kernel loads the associated dynamic loader (we’ll see in a second what that is) and launches it to load the file and do whatever is needed regarding dynamic libraries and symbol resolution. So, when you start a dynamic program, the very first code that is executed is part of the dynamic loader and not part of the binary.

Next stop will be the ELF sections.

ELF Sections

The next interesting flag is -S (not the capital) that will show all the details about the program sections. The output for our example is this:

$ readelf -S challenge01.x86_64.bin
There are 31 section headers, starting at offset 0x37b8:

Section Headers:
  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [ 0]                   NULL             0000000000000000  00000000
       0000000000000000  0000000000000000           0     0     0
  [ 1] .note.gnu.pr[...] NOTE             0000000000000350  00000350
       0000000000000020  0000000000000000   A       0     0     8
  [ 2] .note.gnu.bu[...] NOTE             0000000000000370  00000370
       0000000000000024  0000000000000000   A       0     0     4
  [ 3] .interp           PROGBITS         0000000000000394  00000394
       000000000000001c  0000000000000000   A       0     0     1
  [ 4] .gnu.hash         GNU_HASH         00000000000003b0  000003b0
       0000000000000028  0000000000000000   A       5     0     8
  [ 5] .dynsym           DYNSYM           00000000000003d8  000003d8
       0000000000000120  0000000000000018   A       6     1     8
  [ 6] .dynstr           STRTAB           00000000000004f8  000004f8
       00000000000000af  0000000000000000   A       0     0     1
  [ 7] .gnu.version      VERSYM           00000000000005a8  000005a8
       0000000000000018  0000000000000002   A       5     0     2
  [ 8] .gnu.version_r    VERNEED          00000000000005c0  000005c0
       0000000000000030  0000000000000000   A       6     1     8
  [ 9] .rela.dyn         RELA             00000000000005f0  000005f0
       00000000000000f0  0000000000000018   A       5     0     8
  [10] .rela.plt         RELA             00000000000006e0  000006e0
       0000000000000078  0000000000000018  AI       5    24     8
  [11] .init             PROGBITS         0000000000001000  00001000
       0000000000000017  0000000000000000  AX       0     0     4
  [12] .plt              PROGBITS         0000000000001020  00001020
       0000000000000060  0000000000000010  AX       0     0     16
  [13] .plt.got          PROGBITS         0000000000001080  00001080
       0000000000000008  0000000000000008  AX       0     0     8
  [14] .text             PROGBITS         0000000000001090  00001090
       0000000000000197  0000000000000000  AX       0     0     16
  [15] .fini             PROGBITS         0000000000001228  00001228
       0000000000000009  0000000000000000  AX       0     0     4
  [16] .rodata           PROGBITS         0000000000002000  00002000
       0000000000000075  0000000000000000   A       0     0     8
  [17] .eh_frame_hdr     PROGBITS         0000000000002078  00002078
       000000000000002c  0000000000000000   A       0     0     4
  [18] .eh_frame         PROGBITS         00000000000020a8  000020a8
       00000000000000ac  0000000000000000   A       0     0     8
  [19] .note.ABI-tag     NOTE             0000000000002154  00002154
       0000000000000020  0000000000000000   A       0     0     4
  [20] .init_array       INIT_ARRAY       0000000000003dd0  00002dd0
       0000000000000008  0000000000000008  WA       0     0     8
  [21] .fini_array       FINI_ARRAY       0000000000003dd8  00002dd8
       0000000000000008  0000000000000008  WA       0     0     8
  [22] .dynamic          DYNAMIC          0000000000003de0  00002de0
       00000000000001e0  0000000000000010  WA       6     0     8
  [23] .got              PROGBITS         0000000000003fc0  00002fc0
       0000000000000028  0000000000000008  WA       0     0     8
  [24] .got.plt          PROGBITS         0000000000003fe8  00002fe8
       0000000000000040  0000000000000008  WA       0     0     8
  [25] .data             PROGBITS         0000000000004028  00003028
       0000000000000018  0000000000000000  WA       0     0     8
  [26] .bss              NOBITS           0000000000004040  00003040
       0000000000000010  0000000000000000  WA       0     0     16
  [27] .comment          PROGBITS         0000000000000000  00003040
       000000000000001f  0000000000000001  MS       0     0     1
  [28] .symtab           SYMTAB           0000000000000000  00003060
       00000000000003f0  0000000000000018          29    19     8
  [29] .strtab           STRTAB           0000000000000000  00003450
       000000000000024a  0000000000000000           0     0     1
  [30] .shstrtab         STRTAB           0000000000000000  0000369a
       000000000000011a  0000000000000000           0     0     1
Key to Flags:
  W (write), A (alloc), X (execute), M (merge), S (strings), I (info),
  L (link order), O (extra OS processing required), G (group), T (TLS),
  C (compressed), x (unknown), o (OS specific), E (exclude),
  D (mbind), l (large), p (processor specific)

For each section in the program, it shows the following information:

  • Name. This is just a string, a name for the section. Most sections in C programs have fixed and well-known names. For example, the .text section contains the code of the program, and the .rodata contains read-only data.
  • Type: There are multiple section types. In general, PROGBITS are the most interesting, as they contain information that’s stored in the file.
  • Address: This is the memory address or offset where the section will be located in memory.
  • Offset: This is the offset in the file where the data associated with this section is stored. That file content will end up in memory.
  • Size: This is the size in bytes of the section.
  • EntSize: This only applies for some special sections and indicates the size of the entities it contains. For example, the section init_array contains the array of constructors to be executed before main. Those are pointers, so its entity size is 8 for a 64-bit binary, like the one we’re looking at.
  • Flags: These are flags. Among other things, they contain the intended permissions. As we’ll see in the next section, this is kind of an indication.
  • Link, Info, and Align. Are of little use for us, because the information we can access with them is already shown by readelf using other flags.

The sections themselves are of little utility during reverse engineering. They are mainly intended to help in the linking process but are not needed for the execution of the program, and for that reason, the whole section table is usually removed from malware or other programs as part of the stripping process that also removes the symbols and debug information. If they are available, it may make things a little bit easier but not much.

The structure that is really important are the Program Headers

ELF Program Headers

The ELF Program Headers are the ones really defining how the process will look in memory. Sections have to be mapped to a, let’s say, compatible program header in normal conditions. You’ll see that the information in the file is pretty similar to the one stored for the sections, but the main difference is that this is really relevant.

We can list the program headers in an ELF binary using the -l flag with readelf.

$ readelf -l challenge01.x86_64.bin

Elf file type is DYN (Position-Independent Executable file)
Entry point 0x1090
There are 14 program headers, starting at offset 64

Program Headers:
  Type           Offset             VirtAddr           PhysAddr
                 FileSiz            MemSiz              Flags  Align
  PHDR           0x0000000000000040 0x0000000000000040 0x0000000000000040
                 0x0000000000000310 0x0000000000000310  R      0x8
  INTERP         0x0000000000000394 0x0000000000000394 0x0000000000000394
                 0x000000000000001c 0x000000000000001c  R      0x1
      [Requesting program interpreter: /lib64/ld-linux-x86-64.so.2]
  LOAD           0x0000000000000000 0x0000000000000000 0x0000000000000000
                 0x0000000000000758 0x0000000000000758  R      0x1000
  LOAD           0x0000000000001000 0x0000000000001000 0x0000000000001000
                 0x0000000000000231 0x0000000000000231  R E    0x1000
  LOAD           0x0000000000002000 0x0000000000002000 0x0000000000002000
                 0x0000000000000174 0x0000000000000174  R      0x1000
  LOAD           0x0000000000002dd0 0x0000000000003dd0 0x0000000000003dd0
                 0x0000000000000270 0x0000000000000280  RW     0x1000
  DYNAMIC        0x0000000000002de0 0x0000000000003de0 0x0000000000003de0
                 0x00000000000001e0 0x00000000000001e0  RW     0x8
  NOTE           0x0000000000000350 0x0000000000000350 0x0000000000000350
                 0x0000000000000020 0x0000000000000020  R      0x8
  NOTE           0x0000000000000370 0x0000000000000370 0x0000000000000370
                 0x0000000000000024 0x0000000000000024  R      0x4
  NOTE           0x0000000000002154 0x0000000000002154 0x0000000000002154
                 0x0000000000000020 0x0000000000000020  R      0x4
  GNU_PROPERTY   0x0000000000000350 0x0000000000000350 0x0000000000000350
                 0x0000000000000020 0x0000000000000020  R      0x8
  GNU_EH_FRAME   0x0000000000002078 0x0000000000002078 0x0000000000002078
                 0x000000000000002c 0x000000000000002c  R      0x4
  GNU_STACK      0x0000000000000000 0x0000000000000000 0x0000000000000000
                 0x0000000000000000 0x0000000000000000  RW     0x10
  GNU_RELRO      0x0000000000002dd0 0x0000000000003dd0 0x0000000000003dd0
                 0x0000000000000230 0x0000000000000230  R      0x1

 Section to Segment mapping:
  Segment Sections...
   00
   01     .interp
   02     .note.gnu.property .note.gnu.build-id .interp .gnu.hash .dynsym .dynstr .gnu.version .gnu.version_r .rela.dyn .rela.plt
   03     .init .plt .plt.got .text .fini
   04     .rodata .eh_frame_hdr .eh_frame .note.ABI-tag
   05     .init_array .fini_array .dynamic .got .got.plt .data .bss
   06     .dynamic
   07     .note.gnu.property
   08     .note.gnu.build-id
   09     .note.ABI-tag
   10     .note.gnu.property
   11     .eh_frame_hdr
   12
   13     .init_array .fini_array .dynamic .got

Note how readelf shows the mapping of all the sections to each of the program headers at the very end of the output. Here, there’s more interesting information to obtain. Let’s take a quick look at the different Program Header fields:

  • Type: The most interesting type for us is LOAD that means, whatever is in this offset in the file has to end up in this address in memory (offset and address are other fields that we explain below).
  • Offset: This is the offset in disk where the information associated with this segment is stored.
  • VirtAddr and PhyAddr: These are the memory address where the Program Header will be mapped. For regular binaries both values are the same. They may be different for example on embedded systems where the virtual (the addresses managed by the linker) and the physical (the address in the real device) are different. See Embedded Systems and Linker Scipts.
  • FileSiz and MemSiz: These fields contain the size of the program header in the file and the size it’ll have in memory. They should usually be the same.
  • Flags: This contains the permissions associated with the memory block where the program header will end up. In general, you should only see the following values: R for read-only, RW for read/write, usually segments containing data, and RE read/execute for segments containing code. A segment with permission RWE is a red flag.
  • Align: This indicates the alignment for the program header.

This is really all the information we get and we need, however it’s worth taking a look at how these ELF structures can be abused by malware and how the structures may be affected for those abuses so we can quickly identify patterns that may make our analysis much easier.

NOTE Program Headers and code Injection

Let’s learn a bit more about those program headers starting with the ones of type NOTE, because we already know a lot about notes. So these are the parts of the file containing the information about. In the previous chapter we’ve seen how to extract that information using different tools as well as extracting the data directly from the file. In general, that’s it about notes, however these segments have been abused in the past for easy code injection. To understand this we need to get a better image picture of the structure of an ELF File. Next figure shows a very basic sketch using the information we already have

0x00000000 +-------------+
           | HEADER      | 64 bytes 
0x00000064 +-------------+
           | PHDR        | 56 x Num of PHDR
           +-------------+ 
           | More Info   |
           | Code        |
           | Data        |
           +-------------+
           | Sections    |
           +-------------+

This is roughly how an ELF looks like. So now let’s imagine that some malware wants to inject code into the program. Let’s consider a regular virus that has to copy itself inside the ELF file. The obvious choice is to add the extra code at the very end of the file (we’ll discuss other options later). Let’s see how the ELF will look like, but this time we will zoom a bit into the PHDR table.

0x00000000 +-------------+
           | HEADER      | 64 bytes 
0x00000064 +-------------+
           | PHDR        | 56 x Num of PHDR
           | PDHDR       |
           | INTERP      |
           | LOAD        | 
           | LOAD (RE)   | ----> Offset 0x1000 size 0x231
           | LOAD (R)    | ----> Offset 0x2000 Size 0x174
           | LOAD (RW)   | ----> Offset 0x2dd0 Size 0x280
           | ...         |
0x00001000 +-------------+ 
           | CODE        |
           |   ...       |
0x00002000 +-------------+             
           | roData      |
           +-------------+
           ~             ~
           | Sections    |
EndOfFile  +-------------+
           | INJECTED    |
           | CODE        |
           +-------------+

In order to get the injected code loaded in memory so it can be executed later, it has to be part of a Program Header with read and execution permissions. The only header we have with those permissions is the one containing the code that is at 0x1000. We could patch its size to make it cover the whole file but then we will end up with the whole file in the data segment… which may be an option but is really a poor solution.

There are two correct ways to do this:

  • The first one consists of making the code segment bigger so the new injected code can be fit there. That will change the complete file layout as now all the rest of segments in the file need to be patched to account for the extra size in the code segment, which is one of the first in the PHDR table.
  • The second is to insert a new entry in the PHDR table, which also forces patching almost every ELF structure that comes after that.

Because of these limitations a technique that was developed consisted in reusing one of the NOTE program header entries, that are not really needed for anything to add our new entry pointing to the injected code. As the entry is already and the extra code is being added at the end, there no offset affected and nothing to patch. This technique is pretty straightforward but also pretty obvious.

The binary will have an extra LOAD segment separated from the other ones, and also that segment will have execution permissions which will make such a code injection technique pretty easy to detect.

LOAD Program Headers and code Injection

Let’s look now for a second to the LOAD segments. These are the contents for our test program:

  LOAD           0x0000000000000000 0x0000000000000000 0x0000000000000000
                 0x0000000000000758 0x0000000000000758  R      0x1000
  LOAD           0x0000000000001000 0x0000000000001000 0x0000000000001000
                 0x0000000000000231 0x0000000000000231  R E    0x1000
  LOAD           0x0000000000002000 0x0000000000002000 0x0000000000002000
                 0x0000000000000174 0x0000000000000174  R      0x1000
  LOAD           0x0000000000002dd0 0x0000000000003dd0 0x0000000000003dd0
                 0x0000000000000270 0x0000000000000280  RW     0x1000

Let’s look to the numbers.

The very first segment contains the first 0x758 bytes of the file. Those are directly loaded into memory at the address selected by the kernel. What this actually means is that the ELF header and the Program Header table will be located at the base address of our program, which may be useful.

Then we have the program header to hold the code (read and execution permission). This starts at offset 0x1000 in the file and contains 0x231 bytes.

After that we have the readonly data that lives at offset 0x2000 and has a size of 0x174 bytes. Finally we have the data section living at offset 0x2dd0 in the file with a size of 0x270.

Let’s draw this map

0x0000 -> +--------+ ------------> 0x0000 +--------+
          |HDR     |                      |HDR     |
          |PHDR    | 0x758                |PHDR    | 0x758
0x0758 -> |        |                      |        |
          ~        ~                      ~        ~
0x1000 -> +--------+ ------------> 0x1000 +--------+
          | CODE   | 0x231                | CODE   | 0x231
0x1231    |        |                      |        |
          ~        ~                      ~        ~
0x2000 -> +--------+ ------------> 0x2000 +--------+
          |rodata  | 0x174                | rodata | 0x174
0x2174 -> |        |                      |        | 
          ~        ~                      ~        ~
0x2dd0 -> +--------+ ------------> 0x3dd0 +--------+
          | data   | 0x270                |        |
0x3040    |        |                      |        | 
          ~        ~                      ~        ~

Have you notice it?. Yes, there are holes between the segments. Let’s take a look:

$ xxd -s $((0x758)) -l 128 challenge01.x86_64.bin
00000758: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000768: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000778: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000788: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000798: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000007a8: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000007b8: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000007c8: 0000 0000 0000 0000 0000 0000 0000 0000  ................
$ xxd -s $((0x1231)) -l 128 challenge01.x86_64.bin
00001231: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00001241: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00001251: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00001261: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00001271: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00001281: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00001291: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000012a1: 0000 0000 0000 0000 0000 0000 0000 0000  ................
$ xxd -s $((0x2174)) -l 128 challenge01.x86_64.bin
00002174: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00002184: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00002194: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000021a4: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000021b4: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000021c4: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000021d4: 0000 0000 0000 0000 0000 0000 0000 0000  ................
000021e4: 0000 0000 0000 0000 0000 0000 0000 0000  ................

Those holes between segments are empty, and are usually known as caves. The cave between the code segment and the read-only data segment is called a code cave, because anything inserted in that area will get loaded into the code memory that will have execution permissions.

As you can see, this is another technique to inject code in an existing binary much easier than the ones discussed before. However, the drawback of this technique is that the code caves may or may not exist and you never know its size so, for example a virus, may fit in the code cave in a binary but won’t for another binary, because the size of these caves depends on the amount of code in the program.

There is a pretty curious fact about these code caves. They end up in memory no matter what the size of the associated program header says. That has an interesting consequence. The code will be loaded in memory, but any program relying on the program headers to determine where the code is, as for example objdump or gdb won’t be able to see that code. Actually, gdb won’t show you initially, but if you step through the code and end up in that code (or you figure out the address where it is) you can actually see it.

The DYNAMIC Program Header

Obviously, the DYNAMIC program header is only available on dynamically linked programs but it deserves a mention as it contains quite some important information. Let’s start taking a look to the output from readelf to check what we can find there:

$ readelf -d challenge01.x86_64.bin

Dynamic section at offset 0x2de0 contains 26 entries:
  Tag        Type                         Name/Value
 0x0000000000000001 (NEEDED)             Shared library: [libc.so.6]
 0x000000000000000c (INIT)               0x1000
 0x000000000000000d (FINI)               0x1228
 0x0000000000000019 (INIT_ARRAY)         0x3dd0
 0x000000000000001b (INIT_ARRAYSZ)       8 (bytes)
 0x000000000000001a (FINI_ARRAY)         0x3dd8
 0x000000000000001c (FINI_ARRAYSZ)       8 (bytes)
 0x000000006ffffef5 (GNU_HASH)           0x3b0
 0x0000000000000005 (STRTAB)             0x4f8
 0x0000000000000006 (SYMTAB)             0x3d8
 0x000000000000000a (STRSZ)              175 (bytes)
 0x000000000000000b (SYMENT)             24 (bytes)
 0x0000000000000015 (DEBUG)              0x0
 0x0000000000000003 (PLTGOT)             0x3fe8
 0x0000000000000002 (PLTRELSZ)           120 (bytes)
 0x0000000000000014 (PLTREL)             RELA
 0x0000000000000017 (JMPREL)             0x6e0
 0x0000000000000007 (RELA)               0x5f0
 0x0000000000000008 (RELASZ)             240 (bytes)
 0x0000000000000009 (RELAENT)            24 (bytes)
 0x000000006ffffffb (FLAGS_1)            Flags: PIE
 0x000000006ffffffe (VERNEED)            0x5c0
 0x000000006fffffff (VERNEEDNUM)         1
 0x000000006ffffff0 (VERSYM)             0x5a8
 0x000000006ffffff9 (RELACOUNT)          4
 0x0000000000000000 (NULL)               0x0

I won’t go through all the entries but I will highlight a few of them.

The first one is the NEEDED entries. These are the libraries this binary needs to work. The example below has a single NEEDED library, but you will see entries like this for each library reported by ldd. It is the same information.

In addition to this, the section tells us about the whereabouts of the relocation data, the GOT and PLT table and also about the symbols. This is all information that the dynamic linker needs to load the program in memory and prepare it for execution. Overall, from the reversing point of view, this information is not that useful, except maybe for the location of the GOT that may be manipulated for certain kinds of function bypassing or exploits.

The last field to mention is FLAGS_1. Usually, for binaries, you will just see if the file is a PIE (Position Independent executable), or the flag NOW, meaning that all symbols will be resolved before starting program execution. Anything else should be, at least, checked. Note that for dynamic libraries you may see other flags as many of them are related to the dynamic linking process.

Constructors and Destructors

The DYNAMIC section also gives us information about the constructors and destructors in this binary. Those are functions that get executed before main (constructors), or after returning from main when the program execution finishes (destructors). A value to pay attention to is the size. By default it is 8 (for a 64-bit binary, 4 for a 32-bit one).

The INIT_ARRAY usually contains a pointer to frame_dummy historically related to exception frame information and other runtime data, however in modern systems it is just a stub, but GCC continues emitting it. The FINI_ARRAY usually contains a pointer to __do_global_dtors_aux that is a generic clean-up function which doesn’t do much.

However, if we see extra values in the size, that means the program has constructors or destructors that would get executed automatically at startup or shutdown and won’t be referenced directly in the code. The C run-time is the one that will run those functions when required. This is common, for example with crypters or profiling tools.

To check what is there we can use xxd, but first we need to figure out the offset in the file. Let’s get back for a second to the last LOAD program header in our program:

LOAD           0x0000000000002dd0 0x0000000000003dd0 0x0000000000003dd0
               0x0000000000000270 0x0000000000000280  RW     0x1000

The INIT_ARRAY according to the DYNAMIC section in the file is located at 0x3dd0 that is the initial address where this segment is loaded, however, the offset in the disk is 0x2dd0, so that is the one we have to use:

$ xxd -s $((0x2dd0)) -e -g 8 -l 8 challenge01.x86_64.bin
00002dd0: 0000000000001170                   p.......
$ xxd -s $((0x2dd8)) -e -g 8 -l 8 challenge01.x86_64.bin
00002dd8: 0000000000001130                   0.......

So, our constructor is at 0x1170 and our destructor is at 0x1130. There are different ways to figure out which symbol is there. For doing that we can use different tools but the, let’s say, correct one for this task is nm. This is a very old command (you can tell because it is a two-letter one) intended to print the names of the symbols in a binary, which allegedly is where its name comes from (print NaMes).

$ nm challenge01.x86_64.bin | egrep  "1170|1130"
0000000000001130 t __do_global_dtors_aux
0000000000001170 t frame_dummy

There you go, the two symbols we have been talking about. We can get the pointers in the init_array and fini_array directly from the binary without checking the Program headers, using the objdump tool, but as we will talk about it later I preferred to show you an alternative way to get to the same information. An alternative way is using readelf:

$ readelf -x  .init_array challenge01.x86_64.bin

Hex dump of section '.init_array':
  0x00003dd0 70110000 00000000                   p.......

As a final comment, keep in mind that with stripped binaries all this becomes more difficult. We can extract the pointers the same way, but we won’t get any symbol name so we may have to look deeper. However, if your binary has a single entry in this table you can assume what we have just discussed and only do the deeper analysis if something doesn’t match later.

Symbols

We have just learned about the nm command that allows us to dump all symbols in a binary. We can also do that with readelf using the flag -s:

$ readelf -Ws challenge01.x86_64.bin

Symbol table '.dynsym' contains 12 entries:
   Num:    Value          Size Type    Bind   Vis      Ndx Name
     0: 0000000000000000     0 NOTYPE  LOCAL  DEFAULT  UND
     1: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND __libc_start_main@GLIBC_2.34 (2)
     2: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND strncmp@GLIBC_2.2.5 (3)
     3: 0000000000000000     0 NOTYPE  WEAK   DEFAULT  UND _ITM_deregisterTMCloneTable
     4: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND puts@GLIBC_2.2.5 (3)
     5: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND strlen@GLIBC_2.2.5 (3)
     6: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND printf@GLIBC_2.2.5 (3)
     7: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND fgets@GLIBC_2.2.5 (3)
     8: 0000000000000000     0 NOTYPE  WEAK   DEFAULT  UND __gmon_start__
     9: 0000000000000000     0 NOTYPE  WEAK   DEFAULT  UND _ITM_registerTMCloneTable
    10: 0000000000000000     0 FUNC    WEAK   DEFAULT  UND __cxa_finalize@GLIBC_2.2.5 (3)
    11: 0000000000004040     8 OBJECT  GLOBAL DEFAULT   26 stdin@GLIBC_2.2.5 (3)

Symbol table '.symtab' contains 42 entries:
   Num:    Value          Size Type    Bind   Vis      Ndx Name
     0: 0000000000000000     0 NOTYPE  LOCAL  DEFAULT  UND
     1: 0000000000000000     0 FILE    LOCAL  DEFAULT  ABS Scrt1.o
     2: 0000000000002154    32 OBJECT  LOCAL  DEFAULT   19 __abi_tag
     3: 0000000000000000     0 FILE    LOCAL  DEFAULT  ABS crtstuff.c
     4: 00000000000010c0     0 FUNC    LOCAL  DEFAULT   14 deregister_tm_clones
     5: 00000000000010f0     0 FUNC    LOCAL  DEFAULT   14 register_tm_clones
     6: 0000000000001130     0 FUNC    LOCAL  DEFAULT   14 __do_global_dtors_aux
 
     --snip--

The flag -W does a wide output. By default, readelf limits its output to 80 columns, so if the information to be shown takes more real estate it has to be shortened using ellipsis. The -W overrides that behavior and will print all the information. For dynamic symbols, names are usually a bit long because the library they come from is appended to the names, increasing its length.

Anyway, readelf gives us some additional information, as for instance, if the symbol is a function, a file, or something else. And for local symbols, it even gives us its location. Check for instance the line for our old good friend __do_global_dtors_aux at the end of the listing.

The first part of the readelf output shows us the dynamic symbols. Those are symbols coming from external libraries and they won’t be removed by strip because the dynamic linker needs to know the names of the functions in order to dynamically link the function. So, we can always know which external functions a dynamic program uses. Note that statically linked programs won’t have dynamic symbols.

Those dynamic symbols can give us insights on what we may find in the binary. Some classical examples:

A program with a very short list of dynamic symbols but including the dlopen and dlsym symbols is likely trying to hide the function it uses and prevent hooking using LD_PRELOAD. Usually something fishy is going on.

Other functions are also very suspicious. For example, programs using mmap are relatively normal, but if the mmap/mprotect combo comes together, the program is likely to try some kind of run-time execution or modification.

Overall, a regular program will use quite some functions from external libraries and in those cases, a short list is more suspicious than a long list.

Finally, the actual list of symbols and its memory addresses, if available, is pretty useful for the static analysis of the program. Additionally, sometimes, believe it or not, symbols are not just kept but they are not obfuscated at all giving us even more information. For example, in the Challenge1 symbol dump we can see some enlightening symbols:

Symbol table '.symtab' contains 42 entries:
   Num:    Value          Size Type    Bind   Vis      Ndx Name
     0: 0000000000000000     0 NOTYPE  LOCAL  DEFAULT  UND
     1: 0000000000000000     0 FILE    LOCAL  DEFAULT  ABS Scrt1.o
     (...)
    12: 0000000000004038     8 OBJECT  LOCAL  DEFAULT   25 master_key
    20: 0000000000000000     0 FUNC    GLOBAL DEFAULT  UND strncmp@GLIBC_2.2.5

Yes, the master_key is just a global symbol at position 0x4038 and we can get its content like this:

$ xxd -s $((0x3038)) -e -g 8 -l 8  challenge01.x86_64.bin
00003038: 0000000000002008                   . ......

The master_key symbol is a pointer to a string in the .rodata section. As this is an Intel machine we use the -e to read the value as a Little Endian, and read a pointer that is 8 bytes. Remember that the data program header is mapped at 0x3dd0 but it is located in the disk at 0x2dd0 so address 0x4038 in memory actually maps to 0x3038 in disk.

The master_key points to 0x2008 so we can just dump that address directly from the file.

$ xxd -s $((0x2008))  -l 16  challenge01.x86_64.bin
00002008: 5375 7065 7253 6563 7265 7400 0000 0000  SuperSecret.....

Also note that, the same way symbol names can give us a lot of information, they can also be used to misinform us, suggesting wrong ways to pursue. Overall, in real-world we won’t find unstripped binaries very often so the use of symbols is limited, however it is always good to know how to use them in case it’s needed.

maca

maca is a small program I wrote that mimics readelf but has a slightly different output which I found more useful for me. It doesn’t do everything readelf can do, but it does most of what you would need when reverse engineering a binary. I’ll just show you the same flags we’ve seen for readelf just for illustration purposes. The information shown is the same so no further explanation will be provided.

First, let’s look at the ELF header:

$ maca -h challenge01.x86_64.bin
ELF Header
Description      : [64 bits] [AMD x86-64] [LSB] [System V]  | ABI Version [0] [ET_DYN : Shared object]
ENTRY            : 0x1090
Header Size      :  64  |  0x40
Program Headers  :  14  @  0x40   (56 bytes/entry)
Sections         :  31  @  0x37b8 (64 bytes/entry)
Sections Names   : [30] -> 0x3f38

This is roughly the same information readelf provides but presented in a terser way. Next, we can take a look at the sections.

$ maca -S challenge01.x86_64.bin
SECTIONS
 N     NAME                  TYPE   ADD    OFF     SZ ES   F  L  I   E
[00]                         0        0      0      0 00      0  0   0.000
[01]   .note.gnu.property    7      350    350     20 00   A  0  0   25.688
[02]   .note.gnu.build-id    7      370    370     24 00   A  0  0   54.034
[03]              .interp    1      394    394     1c 00   A  0  0   52.864
[04]            .gnu.hash fff6      3b0    3b0     28 00   A  5  0   34.402
[05]              .dynsym    b      3d8    3d8    120 24   A  6  1   9.835
[06]              .dynstr    3      4f8    4f8     af 00   A  0  0   58.206
[07]         .gnu.version ffff      5a8    5a8     18 02   A  5  0   23.964
[08]       .gnu.version_r fffe      5c0    5c0     30 00   A  6  1   29.253
[09]            .rela.dyn    4      5f0    5f0     f0 24   A  5  0   17.993
[10]            .rela.plt    4      6e0    6e0     78 24  AI  5 24   13.555
[11]                .init    1     1000   1000     17 00  AX  0  0   50.524
[12]                 .plt    1     1020   1020     60 16  AX  0  0   42.734
[13]             .plt.got    1     1080   1080      8 08  AX  0  0   39.865
[14]                .text    1     1090   1090    197 00  AX  0  0   64.240
[15]                .fini    1     1228   1228      9 00  AX  0  0   41.701
[16]              .rodata    1     2000   2000     75 00   A  0  0   58.800
[17]        .eh_frame_hdr    1     2078   2078     2c 00   A  0  0   41.502
[18]            .eh_frame    1     20a8   20a8     ac 00   A  0  0   45.500
[19]        .note.ABI-tag    7     2154   2154     20 00   A  0  0   21.341
[20]          .init_array    e     3dd0   2dd0      8 08  WA  0  0   15.386
[21]          .fini_array    f     3dd8   2dd8      8 08  WA  0  0   13.827
[22]             .dynamic    6     3de0   2de0    1e0 16  WA  6  0   18.228
[23]                 .got    1     3fc0   2fc0     28 08  WA  0  0   0.278
[24]             .got.plt    1     3fe8   2fe8     40 08  WA  0  0   14.621
[25]                .data    1     4028   3028     18 00  WA  0  0   12.710
[26]                 .bss    8     4040   3040     10 00  WA  0  0   50.697
[27]             .comment    1        0   3040     1f 01  MS  0  0   52.909
[28]              .symtab    2        0   3060    3f0 24     29 19   22.136
[29]              .strtab    3        0   3450    24a 00      0  0   61.897
[30]            .shstrtab    3        0   369a    11a 00      0  0   53.544

Again, it shows the same information than readelf with the exception of the last column that represents the entropy of that section (see chapter 1 for details). It is represented as a percentage where 100% is maximal entropy (8 in this case, a complete byte).

The program headers are also shown in a more compact format and also shows the entropy calculated for each of them.

$ maca -l challenge01.x86_64.bin
PROGRAM HEADERS
[  ]             TYPE PERM    VADDR     PADDR       OFFSET FILESIZE  MEMSIZE    ALIGN   ENT
[00]             PHDR [4] R         40       40       40 00000310      310        8     20.900
[01]           INTERP [4] R        394      394      394 0000001c       1c        1     51.297
[02]             LOAD [4] R          0        0        0 00000758      758     1000     32.708
[03]             LOAD [5] R E     1000     1000     1000 00000231      231     1000     63.477
[04]             LOAD [4] R       2000     2000     2000 00000174      174     1000     56.466
[05]             LOAD [6] RW      3dd0     3dd0     2dd0 00000270      280     1000     18.398
[06]          DYNAMIC [6] RW      3de0     3de0     2de0 000001e0      1e0        8     18.171
[07]             NOTE [4] R        350      350      350 00000020       20        8     26.123
[08]             NOTE [4] R        370      370      370 00000024       24        4     54.062
[09]             NOTE [4] R       2154     2154     2154 00000020       20        4     21.931
[10]     GNU_PROPERTY [4] R        350      350      350 00000020       20        8     25.961
[11]     GNU_EH_FRAME [4] R       2078     2078     2078 0000002c       2c        4     39.789
[12]        GNU_STACK [6] RW         0        0        0 00000000        0       10     0.000
[13]         GNU_RELO [4] R       3dd0     3dd0     2dd0 00000230      230        1     17.321

Finally, maca can also show the symbols in the binary. I will show the master_key symbol as we did with readelf so you see that similar information is shown:

$ maca -y challenge01.x86_64.bin
SYMBOLS & DYNAMC SYMBOLS
[   0] - 00000000    0   NOTYPE      UNDEFINED NONAME
[   1] - 00000000    0     FILE            ABS Scrt1.o
(...)
[   6] - 00001130    0     FUNC          .text __do_global_dtors_aux
(..)
[   9] - 00001170    0     FUNC          .text frame_dummy
(...)
[  12] - 00004038    8   OBJECT          .data master_key
[  13] - 00000000    0     FILE            ABS crtstuff.c
(...)

Again you can see that roughly the same information is shown. Finally, maca is also able to show the dynamic section of a file if it is available. It shows the symbols or addresses (if symbols are not available) of the init_array and fini_array tables:

$ maca -d challenge01.x86_64.bin
DYNAMIC SECTION
NEEDED               libc.so.6
INIT                 0x1000
FINI                 0x1228
INIT_ARRAY           0x3dd0 frame_dummy
INIT_ARRAYSZ         0x8
FINI_ARRAY           0x3dd8 __do_global_dtors_aux
FINI_ARRAYSZ         0x8
GNU_HASH             0x3b0
STRTAB               0x4f8
SYMTAB               0x3d8
STRSZ                0xaf
SYMENT               0x18
DEBUG                0x0
PLTGOT               0x3fe8
PLTRELSZ             0x78
PLTREL               0x7
JMPREL               0x6e0
RELA                 0x5f0
RELASZ               0xf0
RELAENT              0x18
FLAGS_1              PIE
VERNEED              0x5c0
VERNEEDNUM           0x1
VERSYM               0x5a8
RELACOUNT            0x4
NULL                 0x0

maca colors its output making it a little more easy to identify where the different elements belong to or if there is something out. For example, it will color red if the constructors or destructors tables have more than one entry. It’ll also highlight some of the functions we mentioned earlier when showing the binary symbols, like dlopen/dlsym or mprotect, as well as well-known ones like main or _start.

readelf summary

readelf is a very powerful tool able to explore all the details of the ELF files and to extract a lot of important information for the next stages of the analysis. readelf can do much more, for example, it can dump debug information and other details. However, from a point of view of reverse engineering, this is not that important because, if the sample already has debug information, that will make everything much easier.

Let’s move on to the next tool that will enable us to start our static analysis of the binary.

objdump

Now things start to become interesting. We’ll start looking inside the files. We’ll start looking at the programs. There are many tools to do this but we will start looking to the simplest one and also the one that you get out of the box in any UNIX system. It’s objdump.

As it happens with readelf, objdump provides a lot of options to look into binaries. We’ll just scratch the surface looking into the options that are more useful for us.

In real life you won’t analyze a binary with objdump. There are much better and powerful tools for doing that, however, using it for some simple examples will help us figure out exactly what those tools do and how they work. Why they can find some information and why other stays hidden or requires special actions to be shown. Consider this as learning the foundations.

objdump allows us to dump many parts from a binary, but mostly the assembly code. It is part of the standard toolchain so you will have different versions for different architectures, however all work the same way. The main flag you will use with objdump is -d that stands for disassembly. This together with a little knowledge of how objdump prints information, converts this program in a powerful inspection tool. The two rules you need to know are:

  • Any recognized symbol is surrounded by angle brackets
  • Symbol position is marked with the name of the symbol between angle brackets followed by a colon.

For example, the next command will show us all the references to main in Challenge01.

$ objdump -d challenge01.x86_64.bin | grep "<main>"
    10a4:        48 8d 3d ce 00 00 00         lea    0xce(%rip),%rdi        # 1179 <main>
0000000000001179 <main>:

The first line is the main reference in the start-up code, while the second one is the actual definition of the main function. We can use the -C, -B and -A grep’s flags to get some more information. -C stands for context and will show us the tell how many lines before and after the match we want to see. -B stands for before and will show us the indicated number of lines before the match, while -A stands for, obviously, after and will show us the lines after the match.

The following command will show us the beginning (first 10 lines) of main:

$ objdump -d challenge01.x86_64.bin | grep -A 10 "<main>:"
0000000000001179 <main>:
    1179:   55                      push   %rbp
    117a:   48 89 e5                mov    %rsp,%rbp
    117d:   48 81 ec 00 04 00 00    sub    $0x400,%rsp
    1184:   48 8d 05 8d 0e 00 00    lea    0xe8d(%rip),%rax        # 2018 <_IO_stdin_used+0x18>
    118b:   48 89 c7                mov    %rax,%rdi
    118e:   e8 ad fe ff ff          call   1040 <puts@plt>
    1193:   48 8d 05 9f 0e 00 00    lea    0xe9f(%rip),%rax        # 2039 <_IO_stdin_used+0x39>
    119a:   48 89 c7                mov    %rax,%rdi
    119d:   e8 9e fe ff ff          call   1040 <puts@plt>
    11a2:   48 8d 05 a2 0e 00 00    lea    0xea2(%rip),%rax        # 204b <_IO_stdin_used+0x4b>

While the following command would help us to identify which function is referencing main.

$ objdump -d challenge01.x86_64.bin | grep -B 10 "<main>$"
0000000000001090 <_start>:
    1090:   31 ed                   xor    %ebp,%ebp
    1092:   49 89 d1                mov    %rdx,%r9
    1095:   5e                      pop    %rsi
    1096:   48 89 e2                mov    %rsp,%rdx
    1099:   48 83 e4 f0             and    $0xfffffffffffffff0,%rsp
    109d:   50                      push   %rax
    109e:   54                      push   %rsp
    109f:   45 31 c0                xor    %r8d,%r8d
    10a2:   31 c9                   xor    %ecx,%ecx
    10a4:   48 8d 3d ce 00 00 00    lea    0xce(%rip),%rdi        # 1179 <main>

Note the use of $ to indicate that we only want to consider the lines where our symbol is at the very end of the line. This way, the main definition is not shown, the same way that adding the colon will match the main definition.

Your first reverse

Before continuing with what objdump could do, let’s focus for a sec on what you can already do and solve the Challenge01. Let’s check its main function.

$ objdump -d challenge01.x86_64.bin | grep -A 40 "<main>:"
0000000000001179 <main>:
    1179:   55                      push   %rbp
    117a:   48 89 e5                mov    %rsp,%rbp
    117d:   48 81 ec 00 04 00 00    sub    $0x400,%rsp
    1184:   48 8d 05 8d 0e 00 00    lea    0xe8d(%rip),%rax        # 2018 <_IO_stdin_used+0x18>
    118b:   48 89 c7                mov    %rax,%rdi
    118e:   e8 ad fe ff ff          call   1040 <puts@plt>
    1193:   48 8d 05 9f 0e 00 00    lea    0xe9f(%rip),%rax        # 2039 <_IO_stdin_used+0x39>
    119a:   48 89 c7                mov    %rax,%rdi
    119d:   e8 9e fe ff ff          call   1040 <puts@plt>
    11a2:   48 8d 05 a2 0e 00 00    lea    0xea2(%rip),%rax        # 204b <_IO_stdin_used+0x4b>
    11a9:   48 89 c7                mov    %rax,%rdi
    11ac:   b8 00 00 00 00          mov    $0x0,%eax
    11b1:   e8 aa fe ff ff          call   1060 <printf@plt>
    11b6:   48 8b 15 83 2e 00 00    mov    0x2e83(%rip),%rdx        # 4040 <stdin@GLIBC_2.2.5>
    11bd:   48 8d 85 00 fc ff ff    lea    -0x400(%rbp),%rax
    11c4:   be 00 04 00 00          mov    $0x400,%esi
    11c9:   48 89 c7                mov    %rax,%rdi
    11cc:   e8 9f fe ff ff          call   1070 <fgets@plt>
    11d1:   48 8b 05 60 2e 00 00    mov    0x2e60(%rip),%rax        # 4038 <master_key>
    11d8:   48 89 c7                mov    %rax,%rdi
    11db:   e8 70 fe ff ff          call   1050 <strlen@plt>
    11e0:   48 89 c2                mov    %rax,%rdx
    11e3:   48 8b 05 4e 2e 00 00    mov    0x2e4e(%rip),%rax        # 4038 <master_key>
    11ea:   48 8d 8d 00 fc ff ff    lea    -0x400(%rbp),%rcx
    11f1:   48 89 ce                mov    %rcx,%rsi
    11f4:   48 89 c7                mov    %rax,%rdi
    11f7:   e8 34 fe ff ff          call   1030 <strncmp@plt>
    11fc:   85 c0                   test   %eax,%eax
    11fe:   75 11                   jne    1211 <main+0x98>
    1200:   48 8d 05 4f 0e 00 00    lea    0xe4f(%rip),%rax        # 2056 <_IO_stdin_used+0x56>
    1207:   48 89 c7                mov    %rax,%rdi
    120a:   e8 31 fe ff ff          call   1040 <puts@plt>
    120f:   eb 0f                   jmp    1220 <main+0xa7>
    1211:   48 8d 05 4e 0e 00 00    lea    0xe4e(%rip),%rax        # 2066 <_IO_stdin_used+0x66>
    1218:   48 89 c7                mov    %rax,%rdi
    121b:   e8 20 fe ff ff          call   1040 <puts@plt>
    1220:   b8 00 00 00 00          mov    $0x0,%eax
    1225:   c9                      leave
    1226:   c3                      ret

This challenge is so basic that you can basically read the C code directly from the ASM. Yes, believe me, you just need to filter out a few things. The first thing you need to know is that objdump tries its best to name memory addresses. It tries to use a symbol if it matches, but otherwise it uses the closest symbol plus an offset.

In the listing above, all those 0x20XX addresses are referenced as different offsets from something called _IO_stdin_used. However, they are something different. Let’s ask our old good friend readelf about the sections in this file:

$ readelf -S challenge01.x86_64.bin
There are 31 section headers, starting at offset 0x37b8:

Section Headers:
  [Nr] Name              Type             Address           Offset
       Size              EntSize          Flags  Link  Info  Align
  [14] .text             PROGBITS         0000000000001090  00001090
       0000000000000197  0000000000000000  AX       0     0     16
  [15] .fini             PROGBITS         0000000000001228  00001228
       0000000000000009  0000000000000000  AX       0     0     4
  [16] .rodata           PROGBITS         0000000000002000  00002000
       0000000000000075  0000000000000000   A       0     0     8
   (...)
  [25] .data             PROGBITS         0000000000004028  00003028
       0000000000000018  0000000000000000  WA       0     0     8

So we can see that code is at addresses 0x1XXX, data is at addresses 0x4XXX and 0x2XXX are addresses in the read only section. That’s constants or literals if you prefer.

Let’s take a look. At this point, we’ve a bunch of tools to do that. We can use xxd, readelf, but we can also use objdump. Let’s use this last one as it is the one we are talking about.

$ objdump -j .rodata -s challenge01.x86_64.bin

challenge01.x86_64.bin:     file format elf64-x86-64

Contents of section .rodata:
 2000 01000200 00000000 53757065 72536563  ........SuperSec
 2010 72657400 00000000 52657665 72736520  ret.....Reverse
 2020 456e6769 6e656572 696e6720 4368616c  Engineering Chal
 2030 6c656e67 65202331 00286329 2049626f  lenge #1.(c) Ibo
 2040 6c636f64 65203230 32360050 61737377  lcode 2026.Passw
 2050 6f72643a 20004163 63657373 20477261  ord: .Access Gra
 2060 6e746564 21004163 63657373 2044656e  nted!.Access Den
 2070 69656421 00                          ied!.

There you go. Those are all the text strings used by the program. The -s flag allows us to dump the content of a section and the -j flag allows us to indicate which section(s) we’re referring to. Now we can identify all those strings:

  • 0x2008 -> SuperSecret - ??
  • 0x2018 -> Reverse Engineering Challenge #1 - msg1
  • 0x2039 -> (c) Ibolcode 2026 - msg2
  • 0x204b -> Password: - prompt
  • 0x2056 -> Access Granted! - access_granted
  • 0x2066 -> Access Denied - access_denied

Here we had to count the offsets manually. Using readelf or maca makes things a bit easier.

$ maca -s challenge01.x86_64.bin | grep ".rodata"
[           .rodata] 0x002008 [00b]: SuperSecret
[           .rodata] 0x002018 [020]: Reverse Engineering Challenge #1
[           .rodata] 0x002039 [011]: (c) Ibolcode 2026
[           .rodata] 0x00204b [00a]: Password:
[           .rodata] 0x002056 [00f]: Access Granted!
[           .rodata] 0x002066 [00e]: Access Denied!
[         .shstrtab] 0x003748 [007]: .rodata
$ readelf -p .rodata challenge01.x86_64.bin

String dump of section '.rodata':
  [     8]  SuperSecret
  [    18]  Reverse Engineering Challenge #1
  [    39]  (c) Ibolcode 2026
  [    4b]  Password:
  [    56]  Access Granted!
  [    66]  Access Denied!

But it is all the same information. So if we get back to the ASM now, it would be pretty obvious what the program does. Let’s substitute all the .rodata pointers for the names we assigned above.

$ objdump -d challenge01.x86_64.bin | grep -A 40 "<main>:"
0000000000001179 <main>:
    1179:   55                      push   %rbp
    117a:   48 89 e5                mov    %rsp,%rbp
    117d:   48 81 ec 00 04 00 00    sub    $0x400,%rsp
    1184:   48 8d 05 8d 0e 00 00    lea    0xe8d(%rip),%rax        # 2018 <msg1>
    118b:   48 89 c7                mov    %rax,%rdi
    118e:   e8 ad fe ff ff          call   1040 <puts@plt>
    1193:   48 8d 05 9f 0e 00 00    lea    0xe9f(%rip),%rax        # 2039 <msg2>
    119a:   48 89 c7                mov    %rax,%rdi
    119d:   e8 9e fe ff ff          call   1040 <puts@plt>
    11a2:   48 8d 05 a2 0e 00 00    lea    0xea2(%rip),%rax        # 204b <prompt>
    11a9:   48 89 c7                mov    %rax,%rdi
    11ac:   b8 00 00 00 00          mov    $0x0,%eax
    11b1:   e8 aa fe ff ff          call   1060 <printf@plt>
    11b6:   48 8b 15 83 2e 00 00    mov    0x2e83(%rip),%rdx        # 4040 <stdin@GLIBC_2.2.5>
    11bd:   48 8d 85 00 fc ff ff    lea    -0x400(%rbp),%rax
    11c4:   be 00 04 00 00          mov    $0x400,%esi
    11c9:   48 89 c7                mov    %rax,%rdi
    11cc:   e8 9f fe ff ff          call   1070 <fgets@plt>
    11d1:   48 8b 05 60 2e 00 00    mov    0x2e60(%rip),%rax        # 4038 <master_key>
    11d8:   48 89 c7                mov    %rax,%rdi
    11db:   e8 70 fe ff ff          call   1050 <strlen@plt>
    11e0:   48 89 c2                mov    %rax,%rdx
    11e3:   48 8b 05 4e 2e 00 00    mov    0x2e4e(%rip),%rax        # 4038 <master_key>
    11ea:   48 8d 8d 00 fc ff ff    lea    -0x400(%rbp),%rcx
    11f1:   48 89 ce                mov    %rcx,%rsi
    11f4:   48 89 c7                mov    %rax,%rdi
    11f7:   e8 34 fe ff ff          call   1030 <strncmp@plt>
    11fc:   85 c0                   test   %eax,%eax
    11fe:   75 11                   jne    1211 <main+0x98>
    1200:   48 8d 05 4f 0e 00 00    lea    0xe4f(%rip),%rax        # 2056 <access_granted>
    1207:   48 89 c7                mov    %rax,%rdi
    120a:   e8 31 fe ff ff          call   1040 <puts@plt>
    120f:   eb 0f                   jmp    1220 <main+0xa7>
    1211:   48 8d 05 4e 0e 00 00    lea    0xe4e(%rip),%rax        # 2066 <access_denied>
    1218:   48 89 c7                mov    %rax,%rdi
    121b:   e8 20 fe ff ff          call   1040 <puts@plt>
    1220:   b8 00 00 00 00          mov    $0x0,%eax
    1225:   c9                      leave
    1226:   c3                      ret

Now you just need to know basic assembler and, in this case, the easiest way to solve it is working backwards. From this point on you really need to know, or at least know how to find out, what simple ASM instructions does and also the standard function conventions. You can find this information from many sources, but my book Heavy Wizardry 101 is a great starting point to get all that foundations.

We need to get to 0x1200 that is the fragment of code that represents getting access to the program.

    11f7:   e8 34 fe ff ff          call   1030 <strncmp@plt>
    11fc:   85 c0                   test   %eax,%eax
    11fe:   75 11                   jne    1211 <main+0x98>
    1200:   48 8d 05 4f 0e 00 00    lea    0xe4f(%rip),%rax        # 2056 <access_granted>
    1207:   48 89 c7                mov    %rax,%rdi
    120a:   e8 31 fe ff ff          call   1040 <puts@plt>
    120f:   eb 0f                   jmp    1220 <main+0xa7>
    1211:   48 8d 05 4e 0e 00 00    lea    0xe4e(%rip),%rax        # 2066 <access_denied>

So, there is a call to strncmp that is a C function to compare two strings (we’ll see in a sec which ones). If the result is not zero, what means that the strings are different, the program will jump to 0x1211 that is the part of the program that denies the access. So we just need to know which two strings are compared.

    11e0:   48 89 c2                mov    %rax,%rdx
    11e3:   48 8b 05 4e 2e 00 00    mov    0x2e4e(%rip),%rax        # 4038 <master_key>
    11ea:   48 8d 8d 00 fc ff ff    lea    -0x400(%rbp),%rcx
    11f1:   48 89 ce                mov    %rcx,%rsi
    11f4:   48 89 c7                mov    %rax,%rdi
    11f7:   e8 34 fe ff ff          call   1030 <strncmp@plt>

The function is taking comparing the string pointed by symbol named master_key, that as per our previous analysis of the process memory map is a global variable, with the string pointed by a local variable created at rbp - 0x400. The content of the local variable is easy to figure out, we just need to check the code just above.

    11b6:   48 8b 15 83 2e 00 00    mov    0x2e83(%rip),%rdx        # 4040 <stdin@GLIBC_2.2.5>
    11bd:   48 8d 85 00 fc ff ff    lea    -0x400(%rbp),%rax
    11c4:   be 00 04 00 00          mov    $0x400,%esi
    11c9:   48 89 c7                mov    %rax,%rdi
    11cc:   e8 9f fe ff ff          call   1070 <fgets@plt>

This is a standard call to fgets where the buffer parameter is actually rbp-0x400. In other words, the program is doing:

fgets (user_input, 0x400, stdin);

So, we can call rbp-0x400 user_input and is the place where the user input will be stored. Now we just need to figure out what is the other strings we’re feeding into strncmp. Let’s take a look to what’s in .data section.

$ objdump -j .data -s challenge01.x86_64.bin

challenge01.x86_64.bin:     file format elf64-x86-64

Contents of section .data:
 4028 00000000 00000000 30400000 00000000  ........0@......
 4038 08200000 00000000                    . ......

Taking into account that we are working with an x86_64 64-bit binary, doing the endianess transform and taking 8 bytes, 0x4038 is a pointer to the address 0x2008 which, according to our previous dump of the .rodata section, is the string SuperSecret.

As you can see, objdump gives you all the information to reverse any program, however, it may be a bit cumbersome and tedious to just use it. We have to go through different steps to extract the different pieces of information to build the whole puzzle. Regular reversing tools do all these steps for us automatically so, for a simple challenge like this, Ghidra, radare2, or IDA will just show you all the strings and the function calls straight away. Anyhow, I always think it is valuable to know what all those tools are actually doing. We’ll see a few more of these stuff later in the book.

So, there you go. Your first reverse!

Dealing with other programming languages

Before finishing this brief introduction to objdump, let’s take a look to Challenge01.5 . This challenge is exactly the same than Challenge01 but implemented using C++, which changes quite some things. Let’s take a look to main.

$ objdump -d challenge01.x86_64.bin | grep -A 30 "<main>:"
00000000000022a9 <main>:
    22a9:   55                      push   %rbp
    22aa:   48 89 e5                mov    %rsp,%rbp
    22ad:   53                      push   %rbx
    22ae:   48 83 ec 28             sub    $0x28,%rsp
    22b2:   48 8d 45 d0             lea    -0x30(%rbp),%rax
    22b6:   48 89 c7                mov    %rax,%rdi
    22b9:   e8 92 fe ff ff          call   2150 <_ZNSt7__cxx1112basic_stringIcSt11char_traitsIcESaIcEEC1Ev@plt>
    22be:   48 8d 05 43 0d 00 00    lea    0xd43(%rip),%rax        # 3008 <_IO_stdin_used+0x8>
    22c5:   48 89 c6                mov    %rax,%rsi
    22c8:   48 8d 05 31 2e 00 00    lea    0x2e31(%rip),%rax        # 5100 <_ZSt4cout@GLIBCXX_3.4>
    22cf:   48 89 c7                mov    %rax,%rdi
    22d2:   e8 e9 fd ff ff          call   20c0 <_ZStlsISt11char_traitsIcEERSt13basic_ostreamIcT_ES5_PKc@plt>
    22d7:   48 8b 15 d2 2c 00 00    mov    0x2cd2(%rip),%rdx        # 4fb0 <_ZSt4endlIcSt11char_traitsIcEERSt13basic_ostreamIT_T0_ES6_@GLIBCXX_3.4>
    22de:   48 89 d6                mov    %rdx,%rsi
    22e1:   48 89 c7                mov    %rax,%rdi
    22e4:   e8 e7 fd ff ff          call   20d0 <_ZNSolsEPFRSoS_E@plt>
    22e9:   48 8d 05 39 0d 00 00    lea    0xd39(%rip),%rax        # 3029 <_IO_stdin_used+0x29>
    22f0:   48 89 c6                mov    %rax,%rsi
    22f3:   48 8d 05 06 2e 00 00    lea    0x2e06(%rip),%rax        # 5100 <_ZSt4cout@GLIBCXX_3.4>
    22fa:   48 89 c7                mov    %rax,%rdi
    22fd:   e8 be fd ff ff          call   20c0 <_ZStlsISt11char_traitsIcEERSt13basic_ostreamIcT_ES5_PKc@plt>
    2302:   48 8b 15 a7 2c 00 00    mov    0x2ca7(%rip),%rdx        # 4fb0 <_ZSt4endlIcSt11char_traitsIcEERSt13basic_ostreamIT_T0_ES6_@GLIBCXX_3.4>
    2309:   48 89 d6                mov    %rdx,%rsi
    230c:   48 89 c7                mov    %rax,%rdi
    230f:   e8 bc fd ff ff          call   20d0 <_ZNSolsEPFRSoS_E@plt>
    2314:   48 8d 05 20 0d 00 00    lea    0xd20(%rip),%rax        # 303b <_IO_stdin_used+0x3b>
    231b:   48 89 c6                mov    %rax,%rsi
    231e:   48 8d 05 db 2d 00 00    lea    0x2ddb(%rip),%rax        # 5100 <_ZSt4cout@GLIBCXX_3.4>
    2325:   48 89 c7                mov    %rax,%rdi
    2328:   e8 93 fd ff ff          call   20c0 <_ZStlsISt11char_traitsIcEERSt13basic_ostreamIcT_ES5_PKc@plt>

Wow… what are all those super long names?. Well, that’s typical C++ (Rust uses a similar naming schema) naming for functions; they are needed for this kind of languages.

Parametric Polymorphism

Object-oriented languages (actually most class-based as there are not many pure object or prototype-based languages out there) usually implement something named parametric polymorphism, that allows the programmer to write different functions, all sharing the same name but each one having its own set of parameters. That is very convenient when writing code making it overall more terse and clear (however, sometimes it may be confusing too).

However, if the programming language is compiled, as it happens with C++ or Rust, the compiler has to emit symbols for each of these functions. Traditionally, the symbol name associated to a function is its function name, but if we have multiple functions with the same name, how can we know which one the symbol refers to? Well, the solution is to mangle the name, which is just adding extra strings to the name based on the parameters types so, each name becomes different as far as their parameters types are different.

If we just need to mangle primitive types, the mangling process is pretty basic. We can add an i for ints and f for floats, etc., but these oriented programming languages may easily accept a parameter that is a reference to an interface and then the process becomes a bit more complicated.

When working at source code level, all this is transparent to the programmer. The compiler will know which function is which and will produce the correct code. But when we look to the result binary, we have to deal with those special names and perform the de-mangling operation on them to make sense of what they mean.

These special names are also used for other things like RTTI (Run-Time Type Identification) and so on. You can ask the compiler to not produce this special name using the extern "C" prefix. This is common in libraries that need to be used from C but also from C++.

Anyhow, don’t despair, there is a tool that will help us make some sense out of such an obfuscator. Let me introduce to you c++filt. This tool just identifies those strange strings and translates them to something closer to what you would see in the source code, which is already cryptic enough for some C++ code.

$ objdump -d challenge01.x86_64.bin | c++filt | grep -A 30 "<main>:"
00000000000022a9 <main>:
    22a9:   55                      push   %rbp
    22aa:   48 89 e5                mov    %rsp,%rbp
    22ad:   53                      push   %rbx
    22ae:   48 83 ec 28             sub    $0x28,%rsp
    22b2:   48 8d 45 d0             lea    -0x30(%rbp),%rax
    22b6:   48 89 c7                mov    %rax,%rdi
    22b9:   e8 92 fe ff ff          call   2150 <std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >::basic_string()@plt>
    22be:   48 8d 05 43 0d 00 00    lea    0xd43(%rip),%rax        # 3008 <_IO_stdin_used+0x8>
    22c5:   48 89 c6                mov    %rax,%rsi
    22c8:   48 8d 05 31 2e 00 00    lea    0x2e31(%rip),%rax        # 5100 <std::cout@GLIBCXX_3.4>
    22cf:   48 89 c7                mov    %rax,%rdi
    22d2:   e8 e9 fd ff ff          call   20c0 <std::basic_ostream<char, std::char_traits<char> >& std::operator<< <std::char_traits<char> >(std::basic_ostream<char, std::char_traits<char> >&, char const*)@plt>
    22d7:   48 8b 15 d2 2c 00 00    mov    0x2cd2(%rip),%rdx        # 4fb0 <std::basic_ostream<char, std::char_traits<char> >& std::endl<char, std::char_traits<char> >(std::basic_ostream<char, std::char_traits<char> >&)@GLIBCXX_3.4>
    22de:   48 89 d6                mov    %rdx,%rsi
    22e1:   48 89 c7                mov    %rax,%rdi
    22e4:   e8 e7 fd ff ff          call   20d0 <std::basic_ostream<char, std::char_traits<char> >::operator<<(std::basic_ostream<char, std::char_traits<char> >& (*)(std::basic_ostream<char, std::char_traits<char> >&))@plt>
    22e9:   48 8d 05 39 0d 00 00    lea    0xd39(%rip),%rax        # 3029 <_IO_stdin_used+0x29>
    22f0:   48 89 c6                mov    %rax,%rsi
    22f3:   48 8d 05 06 2e 00 00    lea    0x2e06(%rip),%rax        # 5100 <std::cout@GLIBCXX_3.4>
    22fa:   48 89 c7                mov    %rax,%rdi
    22fd:   e8 be fd ff ff          call   20c0 <std::basic_ostream<char, std::char_traits<char> >& std::operator<< <std::char_traits<char> >(std::basic_ostream<char, std::char_traits<char> >&, char const*)@plt>
    2302:   48 8b 15 a7 2c 00 00    mov    0x2ca7(%rip),%rdx        # 4fb0 <std::basic_ostream<char, std::char_traits<char> >& std::endl<char, std::char_traits<char> >(std::basic_ostream<char, std::char_traits<char> >&)@GLIBCXX_3.4>
    2309:   48 89 d6                mov    %rdx,%rsi
    230c:   48 89 c7                mov    %rax,%rdi
    230f:   e8 bc fd ff ff          call   20d0 <std::basic_ostream<char, std::char_traits<char> >::operator<<(std::basic_ostream<char, std::char_traits<char> >& (*)(std::basic_ostream<char, std::char_traits<char> >&))@plt>
    2314:   48 8d 05 20 0d 00 00    lea    0xd20(%rip),%rax        # 303b <_IO_stdin_used+0x3b>
    231b:   48 89 c6                mov    %rax,%rsi
    231e:   48 8d 05 db 2d 00 00    lea    0x2ddb(%rip),%rax        # 5100 <std::cout@GLIBCXX_3.4>
    2325:   48 89 c7                mov    %rax,%rdi
    2328:   e8 93 fd ff ff          call   20c0 <std::basic_ostream<char, std::char_traits<char> >& std::operator<< <std::char_traits<char> >(std::basic_ostream<char, std::char_traits<char> >&, char const*)@plt>

Believe it or not, if you know a little bit of C++, this makes more sense. If not, then it still looks like obfuscated code. The code above uses std::cout to print strings in the console, but as you can see, the code is not as straightforward as the puts calls in our C version.

For now, let’s leave this right here. We’ll get back to C++ later as we may have to introduce some language constructs and idioms to start looking into this kind of code with some confidence.

There is also a rustfilt program that does the same thing for Rust binaries.

objdump summary

As it happened with readelf, objdump provides a lot of options, however, from a reverse engineering point of view, the most interesting ones are the ones we used in our basic exercise above. One interesting feature of objdump is that it is able to show the source code together with the ASM. Arguably this is of little to no use in real-world reverse engineering, getting a sample with debug information is something that you won’t ever see… or if you see it is because it doesn’t worth to reverse engineer that thing.

In the simplest case, the sample will be, at least, stripped, and in those cases, objdump will just dump all the ASM without any structure so we’ll need some extra effort to process it.

Said that, I’ll keep using objdump in this book so, whenever we need to do something, we’ll have to figure out how to do it and this way we learn how the real-world tools work. Note that in some cases we may have to develop small tools to get functionalities not implemented in the regular command-line tools. After all, objdump actually does what it says it’s going to do: Dump object files

SUMMARY

In this chapter we’ve introduced some of the most well-known tools for static analysis of binaries. As I said, there are great tools out there that make many of the tasks we’ll be doing in this book trivial, however, I’d like to emphasize again the importance of knowing what you are doing. When you get your foundations right, then you’ll get the most out of the more powerful tools. Soon we’ll start using these tools, and the ones introduced in next chapter to do some real reversing. That will be much more fun.

Related Posts

Reverse Engineering Course. Part I

Reverse Engineering Tools. File Information

Read

RE_TUTORIALRETOOLSCLI

Reverse Engineering Challenge 1

Get started with the simplest challenge

Read

RE_CHALLENGERELAB

Reverse Enginnering

All reverse engineering information in a single place

Read

REVERSE ENGINEERING RE RE CHALLENGES

Return to Home Page