2013年2月27日 星期三

http://www.natural-science.or.jp/laboratory/virtual_laboratory001_introduction.html#1

2013年2月24日 星期日

gnuplot color names

http://www.uni-hamburg.de/Wiss/FB/15/Sustainability/schneider/gnuplot/colors.htm

2013年2月6日 星期三

On Stack Replacement

From this post, it said:
It’s used to convert a running function’s interpreter frame into a JIT’d frame – in the middle of that method. 
[On-Stack-Replacement (OSR) compilation was first introduced in the famous hotspot server paper, to the best of my knowledge.]

A simple example copied from the same post:

public static void main(String args[]) {
    S1; i=0;
    loop:
    if(P) goto done
      S3; A[i++];
    goto loop; // <<--- here="" osr="" span="">
    done:
    S2;
  }

OSR-compiled function:


void OSR_main() {
    A=value on entry from interpreter;
    i=value on entry from interpreter;
    goto loop;
    loop:
    if(P) goto done
      S3; A[i++];
    goto loop;
    done:
    ...never reached...
}

after OSR_main is compiled, the execution will transfer from goto loop in the interpreter to the OSR_main.

But I am not quiet clear about why ``done'' part in OSR_main is never reached?

2013年2月3日 星期日

Interrupt handling

In emulation, some asynchronous events may arrive in the middle of execution.
Two kinds of asynchronous events are interrupts and exception.

Semantics of Interrupts are explained in this paper.

First, it is a hardware-supported asynchronous transfer of control to an interrupt vector based on the signaling of some condition external to the processor core. An interrupt vector is a dedicated or configurable location in memory that specifies the address to which execution should jump when an interrupt occurs. Second, an interrupt is the execution of an interrupt handler : code that is reachable from an interrupt vector. 

It is irrelevant whether the interrupting condition originates on-chip (e.g., timer expiration) or off-chip (e.g., closure of a mechanical switch). Interrupts usually, but not always, return to the flow of control that was interrupted. Typically an interrupt changes the state of main memory and of device registers, but leaves the main processor context (registers, page tables, etc.) of the interrupted computation undisturbed.


2013年1月29日 星期二

IBM Developer Works

IBM Developer Works
http://www.ibm.com/developerworks/

2013年1月27日 星期日

Indirect branch profiling:

1. what is the frequency of 1-target indirect branches in benchmarks?

1-target indirect branch
there is only one target of this indirect branch instruction during execution.
therefore when we say n-target indirect branch, we mean indirect branches that has n targets during execution.

frequency 
Assume there are 100 guest indirect branches executed (dynamic count) during execution, and there are 8 1-target indirect branches executed. The frequency is 8/100 = 8%.

N-Target Indirect Branches Frequency Distribution of 400.perlbench with diffmail.pl test input:

seems not very useful information

2. what is the hit ratio of last-target prediction of indirect branches?

last-target prediction:
Predict next target of one indirect branch with its last target.

400.perlbench with ref. input diffmail.pl:

Overall 73.63%
RET 93.99%
INDIRECT_CALL 20.45%
UNCOND_INDIRECT_JMP 48.98%

gcc with ref. input 166.i:
Overall 64.77%
RET 61.08%
INDIRECT_CALL 97.93%
UNCOND_INDIRECT_JMP 66.35%

gcc with ref. input scilab.i
Overall 58.93%
RET 61.54%
INDIRECT_CALL 92.98%
UNCOND_INDIRECT_JMP 44.43%

445.gobmk with ref. test  input nngs.tst :
Overall 56.17%
RET 56.01%
INDIRECT_CALL 78.30%
UNCOND_INDIRECT_JMP 67.25%

mini cache prediction is not helpful, it degrades performance about 4%.




2013年1月25日 星期五

Some interesting slides/papers about trace optimization in JVM

http://researcher.watson.ibm.com/researcher/files/us-pengwu/challeng-potential-trace-compilation.pdf

http://researcher.watson.ibm.com/researcher/files/us-pengwu/oopsla111-wu.pdf

http://researcher.watson.ibm.com/researcher/files/us-pengwu/UIUC-Seminar-Scripting-Languages-05-03.pdf

SINOF: A dynamic-static combined framework for dynamic binary translation
http://dl.acm.org/citation.cfm?id=2350593&CFID=174278273&CFTOKEN=47796794

Similar to permanent code cache, previous compiled blocks are saved and loaded by future runs.
Saved blocks are analyzed and optimized by runtime profiling information.
1. What kind of analysis and optimization they used?
2. What kind of information do they collect at runtime?
3. What's the benefit?
First, they use their own IR, and explain why they don't use LLVM or UQBT IR.
Second, in their evaluation, both guest ISA and host ISA are IA32! But they achieved on average 1.38X normalized by native execution time.


A low-overhead dynamic optimization framework for multicores
http://dl.acm.org/citation.cfm?id=2370899&CFID=174278273&CFTOKEN=47796794
Don't know what they do from abstract.
A very short paper (2-page), but I still have no idea what they did.

Adaptive multi-level compilation in a trace-based Java JIT compiler



Adaptive multi-level compilation in a trace-based Java JIT compiler


http://dl.acm.org/citation.cfm?id=2384630&CFID=173817186&CFTOKEN=20831145
an extended work of IBM's Trace-based JVM published in CGO 2011.

The ``meat'' of this paper is : Trace Recompilation via Trace-Transition Graph.

That is, they select traces to be recompiled via Trace-Transition Graph.

First, the recompiled code fragments are still TRACE! not region.
Their scenario is there may be short fragmented traces due to the limit of maximum length in the initial trace building phase. They set max-trace-length to two.
They would like to merge those fragmented traces into one trace.

The other interesting part is the construction of the Trace-Transition Graph.
The information needed are : 1. transition between traces, and 2. how frequency between transition.
They did not use hardware performance monitoring information.
Instead, they periodically check which transition between traces and record the frequency.
They use Branch-and-link instruction for transitions between traces, rather than using jump.
In this approach, the linker register record who the source trace is.

2013年1月13日 星期日

work log

TRACE, 16370, 85460d0

2012年12月27日 星期四

Shack = Shadow Stack
IBTC', Shack, Shack + IBTC'
400.perlbench 1275 1315 1323

2^18 V.S 2^11

2012年12月14日 星期五

Ref. Input, GCC
166.i : OK
c-typeck.i : OK
200.i : NG(10m37s)
cp-decl.i : NG, Segfault (11m55s)
expr.i : NG, (14m9s)
expr2.i : NG, (18m18s)
g23.i : 
due to LEApcrel fails to handle offset to constant pool.

2012年12月12日 星期三

0x080b9249
0x080b9249

r6 0x40ffca1c
r3 0x40ffca30

VST1LNd32_UPD
R2 = op R2, 0, R0, D18, 0
f 4 c 2 2 8 0 4
c: D = 1
2: Rn = 2
2: Vd = 16*D+2 = 18
8: size = 0b10:
0: index_align = 0
4: Rm = 4! Wrong!!!!

2012年12月4日 星期二

Work log

lib/Target/ARM/ARMISelDAGToDAG.cpp
include/llvm/CodeGen/ValueTypes.td

2012年12月2日 星期日

ARM SPEC CPU2006 Native Run

ARM SPEC CPU2006
gcc flags:
-static -O3 -marm -march=armv7-a -mtune=cortex-a8 -mcpu=cortex-a8 -mfloat-abi=softfp -mfpu=neon -ffast-math -ftree-vectorize -funroll-all-loops

ARM Native Run with Ref. input

400.perlbench    2878  2879
401.bzip2        5007  4975
403.gcc          2942  2942
429.mcf          4082  4081
445.gobmk        3265  3267
456.hmmer        3268  3280
458.sjeng        3952  3935
462.libquantum  19697 19732
464.h264ref      4700  4718
471.omnetpp      2772  2782
473.astar        3002  2969
483.xalancbmk    2772  2766

2012年11月30日 星期五

Worklog

find equivalent neon instruction for the following SSE instruction:
PCMPEQBrr
 PCMPEQDrr
 PSLLDri
 PSUBUSBrr
 PUNPCKHBWrr
 PUNPCKHWDrr
 PUNPCKLBWrr
 PUNPCKLDQrr
 PUNPCKLQDQrr
 PUNPCKLWDrr

Trace Mode, Ref Input

Trace Mode, CINT2006 reference input:
400.perlbench    9770    6657         1.47  *
401.bzip2        9650    2623               VE
403.gcc          8050    6133               RE
429.mcf          9120       9.01            RE
445.gobmk       10490   10356         1.01  *
456.hmmer        9330    6044         1.54  *
458.sjeng       12100   10927         1.11  *
462.libquantum  20720   22428         0.924 *
464.h264ref     22130   11599         1.91  *
471.omnetpp      6250     204               RE
473.astar        7020    4235         1.66  *
483.xalancbmk    6900    6767         1.02  *

Fail to run 401, 403, 429, 471
429.mcf needs 839 MB buf we only have 913MB in ARM, not enough memory.
After creating a swap of 1G, mcf can run successfully.
in top, only 17MB in SWAP area.

2012年11月29日 星期四

IBTC + Victim Performance Evaluation

     IBTC      IBTC+Victim
2^11 0.997770  0.999741
2^12 0.999387  0.999932
2^13 0.999725  0.999981
2^14 0.999857  0.999992
2^15 0.999887
2^16 0.999949
2^17 0.999955
2^18 0.999995


Block Mode
Performance (No inline)
IBTC(2^11)+Victim
T: 369.863281   0.545624        138.762726      0.000000        230.276672      0.278259
IBTC(2^18)
T: 375.984528   0.372223        138.313660      0.000000        237.024841      0.273804

Performance - Inline
IBTC(2^11)+Victim
T: 367.764587   0.378113        141.984161      0.000000        225.127594      0.274719
IBTC(2^18)
T: 373.440033   0.549896        145.442383      0.000000        227.171478      0.276276
IBTC(2^12)+Victim (make ibtc table 2^16 KB)
T: 4m20s