Repository navigation
Expand file tree
/
Copy pathfp.txt
More file actions
2014 lines (1484 loc) · 70.6 KB
/
Copy pathfp.txt
File metadata and controls
2014 lines (1484 loc) · 70.6 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
== Functional Programming Techniques ==
=== What does functional programming have to do with objects? ===
It doesn't take much time to realize that Ruby is a deeply object oriented
language that openly steals from the best of Smalltalk. But Matz is an
equal opportunity thief, and has snatched features from various other
languages as well, including Lisp. This means that although Ruby has its
roots in object oriented principles, we also have some of the high level
constructs that facilitate functional programming techniques.
Rightfully speaking, Ruby is not a functional programming language. Though
it is possible to come close with great effort, Ruby simply lacks a number
of the key aspects of functional languages. Virtually all state in Ruby is
mutable, and because it is an object oriented language, we tend to focus
on state management rather than elimination of state. Ruby lacks tail
call optimization footnote:[This can be enabled in YARV at compile time, but
is still experimental], making recursion highly inefficient. Beyond these key
things, there are plenty of other subtleties that purists can discuss at
length if you let them.
However, if you let Ruby be Ruby, you can benefit from certain functional
techniques while still writing object oriented code. This chapter will walk
through several practices that are inspired from functional languages that
have practical value for solving problems in Ruby. We'll look at things
like lazy evaluation, memoization, infinite lists, higher order procedures
and other stuff that has a nice academic ring to it. Along the way, I'll do
my best to show you that this stuff isn't simply abstract and mathematical,
but can meaningfully be used in day to day code.
Let's start by taking a look at how we already use lazy evaluation whether
we realize it or not, and then walk through a popular Ruby library that can
simplify this for us. As in other chapters, we'll take a peek under the hood
to see what's really going on.
=== Laziness Can Be A Virtue (A look at lazy.rb) ===
If you've been writing Ruby for any amount of time, you've probably already
written some code that makes use of lazy evaluation. Before we go farther,
let's look at a couple examples of lazy code you will likely recognize on
sight, if not by name.
For starters, +Proc+ objects are, by definition, lazy:
..............................................................................
a = lambda { File.read("foo.txt") }
..............................................................................
When you create a +Proc+ object using +lambda+, the code is not actually executed
until you call the block. Therefore, a +Proc+ is essentially a chunk of code
that gets executed on demand, rather than in place. So if we actually wanted
to read the contents of this file, we'd need to do:
..............................................................................
b = a.call
..............................................................................
In essence, code is said to be evaluated lazily if it is executed only at the
time it is actually needed, not at the time it was defined. However, this
behavior is not necessarily limited to blocks, we can do this with populating
data for our objects as well. Let's take a look at a simplified model of
Prawn's table cells for a more complete example:
..............................................................................
class Cell
FONT_HEIGHT = 10
FONT_WIDTH = 8
def initialize(text)
@text = text
end
attr_accessor :text
attr_writer :width, :height
def width
@width ||= calculate_width
end
def height
@height ||= calculate_height
end
def to_s
"Cell(#{width}x#{height})"
end
private
def calculate_height
@text.lines.count * FONT_HEIGHT
end
def calculate_width
@text.lines.map { |e| e.length }.max * FONT_WIDTH
end
end
..........................................................................
In this example, +Cell#width+ and +Cell#height+ can either be calculated based
on the text in the cell, or they can be manually set. Because we don't need
to know the exact dimensions of a cell until we render it, this is a perfect
case for lazy evaluation. Even though the calculations may not be expensive
on their own, they add up when dealing with thousands or hundreds of
thousands of cells. Luckily, it's easy to avoid any unnecessary work.
By just looking at the core bits of this object, we can get a clearer sense
of what's going on:
..........................................................................
class Cell
attr_writer :width, :height
def width
@width ||= calculate_width
end
def height
@height ||= calculate_height
end
end
..........................................................................
It should now be plain to see what's happening. If +@width+ or +@height+ have
already been set, their calculations are never run. We also can see that
these calculations will in the worst case be run exactly once, storing the
return value as needed.
The idea here is now we won't have to worry about calculating dimensions of
pre-set +Cell+ objects, and that those that calculate their dimensions will not need
to repeat that calculation each time they are used. I am hoping that readers
are familiar with this Ruby idiom already, but if you are struggling with it
a little bit, just stop for a moment and toy with this in irb and it
should quickly become clear how things work.
..........................................................................
>> cell = Cell.new("Chunky Bacon\nIs Very\nDelicious")
>> cell.width = 1000
=> 1000
>> cell.to_s
=> "Cell(1000x30)"
>> cell.height = 500
=> 500
>> cell.to_s
=> "Cell(1000x500)"
..........................................................................
Though this process is relatively straightforward, we can probably make it
better. It's sort of annoying to have to build special accessors for
our +@width+ and +@height+ when what we're ultimately doing is setting default
values for them, which is normally something we do in our constructor. What's
more, our solution feels a little primitive in nature, at least aesthetically.
This is where MenTaLguY's 'lazy.rb' comes in. It provides a method called
+promise()+ that does exactly what we want, in a much nicer way. The following
code can be used to replace our original implementation:
..........................................................................
require "lazy"
class Cell
FONT_HEIGHT = 10
FONT_WIDTH = 8
def initialize(text)
@text = text
@width = promise { calculate_width }
@height = promise { calculate_height }
end
attr_accessor :text, :width, :height
def to_s
"Cell(#{width}x#{height})"
end
private
def calculate_height
@text.lines.count * FONT_HEIGHT
end
def calculate_width
@text.lines.map { |e| e.length }.max * FONT_WIDTH
end
end
..........................................................................
Gone are our special accessors, and in their place, we have simple +promise()+
calls in our constructor. This method returns a simple proxy object that
wraps a block of code that is designed to be executed later. Once you call
any methods on this object, it passes them along to whatever your block
evaluates to. If this is tough to get your head around, again +irb+ is your
friend:
..........................................................................
>> a = promise { 1 }
=> #<Lazy::Promise computation=#<Proc:0x3ce218@(irb):2>>
>> a + 3
=> 4
>> a
=> 1
..........................................................................
+Lazy::Promise+ objects are very cool, but a little bit sneaky. Because they
essentially work the same as the evaluated object once the block is run once,
it's hard to know that you're actually working with a proxy object. But if
we dig, deep, we can find the truth:
..........................................................................
>> a.class
=> Fixnum
>> a.__class__
=> Lazy::Promise
..........................................................................
If we try out the same examples as we did before with +Cell+, you'll see it
works as expected:
..........................................................................
>> cell = Cell.new("Chunky Bacon\nIs Very\nDelicious")
>> cell.width = 1000
=> 1000
>> cell.to_s
=> "Cell(1000x30)"
>> cell.height = 500
=> 500
>> cell.to_s
=> "Cell(1000x500)"
..........................................................................
Seeing that the output is the same, we can be satisfied knowing that our
'lazy.rb' solution will do the trick, while making our code a little cleaner
looking. However, we shouldn't just pass the work +promise()+ is doing off
as magic, it is worth looking at both for enrichment and because it's
genuinely cool code.
I've went ahead and simplified the implementation of +Lazy::Promise+ a bit,
removing some of the secondary features while still preserving the core
functionality. We're going to look at the naive implementation, but please
use this only for studying. The official version from MenTaLguY can handle
thread synchronization and does much better error handling than what you will
see here, so that's what you'll want if you plan to use 'lazy.rb' in real code.
That all having been said, here's how you'd go about implementing a basic
+Promise+ object on Ruby 1.9:
..........................................................................
module NaiveLazy
class Promise < BasicObject
def initialize(&computation)
@computation = computation
end
def __result__
if @computation
@result = @computation.call
@computation = nil
end
@result
end
def inspect
if @computation
"#<NaiveLazy::Promise computation=#{ @computation.inspect }>"
else
@result.inspect
end
end
def respond_to?( message )
message = message.to_sym
[:__result__, :inspect].include?(message) ||
__result__.respond_to? message
end
def method_missing(*a, &b)
__result__.__send__(*a, &b)
end
end
end
..........................................................................
Though compact, this might look a little daunting in one big chunk, so let's
break it down. First, you'll notice that +NaiveLazy::Promise+ inherits from
+BasicObject+ rather than +Object+:
..........................................................................
module NaiveLazy
class Promise < BasicObject
end
end
..........................................................................
+BasicObject+ omits most of the methods you'll find on +Object+, making it so
you don't need to explicitly remove those methods on your own in order to
create a proxy object. This in effect gives you a blank slate object, which
is exactly what we want for our purposes.
The proxy itself works through `method_missing`, handing virtually all messages
to the result of `+++Promise#__result__+++`:
..........................................................................
module NaiveLazy
class Promise < BasicObject
def method_missing(*a, &b)
__result__.__send__(*a, &b)
end
end
end
..........................................................................
For the uninitiated, this code essentially just allows `promise.some_function`
to be interpreted as `+++promise.__result__.some_function+++`, which makes sense
when we recall how things worked in Cell:
..........................................................................
>> cell.width
=> #<Lazy::Promise computation=#<Proc:...>>
>> cell.width + 10
=> 114
..........................................................................
+Lazy::Promise+ knew how to evaluate the computation and then pass your
message to its result. This `method_missing` trick is how it works under
the hood. When we go back and look at how `+++__result__+++` is implemented, this
becomes even more clear:
..........................................................................
module NaiveLazy
class Promise < BasicObject
def initialize(&computation)
@computation = computation
end
def __result__
if @computation
@result = @computation.call
@computation = nil
end
@result
end
def method_missing(*a, &b)
__result__.__send__(*a, &b)
end
end
end
..........................................................................
When we create a new promise, it stores a code block to be executed later.
Then, when you call a method on the promise object, `method_missing` runs the
`+++__result__+++` method. This method checks to see if there is a +Proc+ object in
`@computation` that needs to be evaluated. If there is, it stores the return
value of that code, and then wipes out the +@computation+. Further calls to
`+++__result__+++` return this value immediately.
On top of this, we add a couple methods to make the proxy more well behaved,
so that you can fully treat an evaluated promise as if it were just an
ordinary value.
..........................................................................
module NaiveLazy
class Promise < BasicObject
# ...
def inspect
if @computation
"#<NaiveLazy::Promise computation=#{ @computation.inspect }>s"
else
@result.inspect
end
end
def respond_to?( message )
message = message.to_sym
[:__result__, :inspect].include?(message) or
__result__.respond_to? message
end
# ...
end
end
..........................................................................
I won't go into much detail about this code, both of these methods essentially
just forward everything they can to the evaluated object, and are nothing
more than underplumbing. However, since these round out the full object,
you'll see them in action when we take our new `NaiveLazy::Promise` for a spin:
..........................................................................
>> num = NaiveLazy::Promise.new { 3 }
=> #<NaiveLazy::Promise computation=#<Proc:0x3cfd98@(irb):2>
>> num + 100
=> 103
>> num
=> 3
>> num.respond_to?(:times)
=> true
>> num.class
=> Fixnum
..........................................................................
So what we have here is an implementation of a proxy object that doesn't
produce a exact value until it absolutely has to. This proxy is fully
transparent so that even though in actuality, `num` here is a `NaiveLazy::Promise`
instance, not a `Fixnum`, other objects in your system won't know or care.
This can be pretty handy when delaying your calculations to the last possible
moment is important.
The reason I showed how to implement a promise is to give you a sense of what
the primitive tools in Ruby are capable of doing. We've seen blocks used in
a lot of different ways throughout Ruby, but this particular case might be
easy to overlook. Another interesting factor here is that although this
concept belongs to the functional programming paradigm,it is also easy to
implement using Ruby's object oriented principles.
As we move on to look at other techniques in this chapter, keep this general
idea in mind. Much will be lost in translation if you try to directly
convert functional concepts into Ruby, but by playing to Ruby's strengths,
you can often preserve the idea without things feeling alien.
We're about to look at some other things that are handy to have in your
toolbelt, but before we do that, I'll reiterate some key points about lazy
evaluation in Ruby.
* Lazy evaluation is useful when you have some code that may never need to
be run, or would best be run as late as possible, especially if this code
is expensive computationally. If you do not have this need, it is better
to do without the overhead.
* All code blocks in Ruby are lazy, and are not executed until explicitly
called.
* For simple needs, you can build attribute accessors for your objects that
avoid running a calculation until they are called, storing the result in
an instance variable once it is executed.
* MenTaLguY's 'lazy.rb' provides a more comprehensive solution for lazy
evaluation, which can be made to be thread safe and is generally more
robust than the naive example shown in this chapter.
We'll now move on to the sticky issue of state maintenance and side-effects,
as this is a key aspect of functional programming that cuts a little against
the grain of traditional Ruby code.
=== Minimizing Mutable State and Reducing Side Effects ===
Although Ruby is object oriented, and therefore relies heavily on mutable
state, we can write non-destructive code in Ruby. In fact, many of our
`Enumerable` methods are inspired by this.
For a trivial example, we can consider the use case for `Enumerable#map`. We
could write our own naive map implementation rather easily:
..........................................................................
def naive_map(array)
array.each_with_object([]) { |e, arr| arr << yield(e) }
end
..........................................................................
When we run this code, it has the same results as `Enumerable#map`, as show
below:
..........................................................................
>> a = [1,2,3,4]
=> [1, 2, 3, 4]
>> naive_map(a) { |x| x + 1 }
=> [2, 3, 4, 5]
>> a
=> [1, 2, 3, 4]
..........................................................................
As you can see, a new array is produced, rather than modifying the original
array. In practice, this is how we tend to write side effect free code in
Ruby. We traverse our original data source, and then build up the results
of a state transformation in a new object. In this way, we don't modify the
original object. Because `naive_map()` doesn't make changes to
anything outside of the function, we can say that this code is side-effect
free.
However, this code still uses mutable state to build up its return value.
To truly make the code stateless, we'd need to build a new array
every time we append a value to an array. Notice the difference between
these two ways of adding a new element to the end of an array:
..........................................................................
>> a
=> [1, 2, 3, 4]
>> a = [1,2,3]
=> [1, 2, 3]
>> b = a << 1
=> [1, 2, 3, 1]
>> a
=> [1, 2, 3, 1]
>> c = a + [2]
=> [1, 2, 3, 1, 2]
>> b
=> [1, 2, 3, 1]
>> a
=> [1, 2, 3, 1]
..........................................................................
It turns out that `Array#<<` modifies its receiver, and `Array#+` does not. With
this knowledge, we can re-write `naive_map` to be truly stateless:
..........................................................................
def naive_map(array, &block)
return [] if array.empty?
[ yield(array[0]) ] + naive_map(array[1..-1], &block)
end
..........................................................................
This code works in a different way, building up the result set by calling
itself repeatedly, resulting in something like this:
..........................................................................
[1,2,3,4] => [2], [2,3,4] => [2] + [3], [3,4] => [2] + [3] + [4], [4] =>
[2] + [3] + [4] + [5] => [2,3,4,5]
..........................................................................
Depending on your taste for recursion, you may find this solution beautiful
or scary. In Ruby, recursive solutions may look elegant and have appeal
from a purist's perspective, but when it comes to their drawbacks, the pragmatists
win out. Other languages optimize for this sort of thing, but Ruby does not,
which is made obvious
by this benchmark:
..........................................................................
require "benchmark"
def naive_map(array, &block)
new_array = []
array.each { |e| new_array << block.call(e) }
return new_array
end
def naive_map_recursive(array, &block)
return [] if array.empty?
[ yield(array[0]) ] + naive_map_recursive(array[1..-1], &block)
end
N = 100_000
Benchmark.bmbm do |x|
a = [1,2,3,4,5]
x.report("naive map") do
N.times { naive_map(a) { |x| x + 1 } }
end
x.report("naive map recursive") do
N.times { naive_map_recursive(a) { |x| x + 1 } }
end
end
# Outputs:
sandal:fp $ ruby naive_map_bench.rb
Rehearsal -------------------------------------------------------
naive map 0.370000 0.010000 0.380000 ( 0.373221)
naive map recursive 0.530000 0.000000 0.530000 ( 0.539722)
---------------------------------------------- total: 0.910000sec
user system total real
naive map 0.360000 0.000000 0.360000 ( 0.369269)
naive map recursive 0.530000 0.000000 0.530000 ( 0.538872)
..........................................................................
Even though our functions are somewhat trivial, we see our recursive solution
performing significantly slower than the iterative one. The reason behind
this is the very high cost of method dispatch in Ruby. This means that
despite the identical complexity between our iterative and recursive
solutions, the latter can quickly become a performance nightmare. If we use
a larger data set, we can see this only exacerbates the problem:
..........................................................................
N = 100_000
Benchmark.bmbm do |x|
a = [1,2,3,4,5] * 20
x.report("naive map") do
N.times { naive_map(a) { |x| x + 1 } }
end
x.report("naive map recursive") do
N.times { naive_map_recursive(a) { |x| x + 1 } }
end
end
# output
sandal:fp $ ruby naive_map_bench.rb
Rehearsal -------------------------------------------------------
naive map 4.360000 0.020000 4.380000 ( 4.393069)
naive map recursive 9.420000 0.030000 9.450000 ( 9.498580)
--------------------------------------------- total: 13.830000sec
user system total real
naive map 4.350000 0.010000 4.360000 ( 4.382038)
naive map recursive 9.420000 0.050000 9.470000 ( 9.532602)
..........................................................................
An important thing to remember is that any recursive solution can be
re-written iteratively. We can actually build an iterative, stateless, naive
map without much extra effort:
..........................................................................
def naive_map_via_inject(array, &block)
array.inject([]) { |s,e| [ yield(e) ] + s }
end
..........................................................................
`Enumerable#inject` is a favorite feature among Rubyists for accumulation. The
way that it works is that it passes two objects into the block: The base
object and the current element. After each step through the iteration, the
return value of the block becomes the new base. Essentially, this code is
doing the same thing our recursive code did, without the recursion. We can
take a quick look at the benchmarks now, expecting some improvement by cutting
out all those expensive recursive method calls:
..........................................................................
N = 100_000
Benchmark.bmbm do |x|
a = [1,2,3,4,5] * 20
x.report("naive map") do
N.times { naive_map(a) { |x| x + 1 } }
end
x.report("naive map recursive") do
N.times { naive_map_recursive(a) { |x| x + 1 } }
end
x.report("naive map via inject") do
N.times { naive_map_via_inject(a) { |x| x + 1 } }
end
end
# Output
sandal:fp $ ruby naive_map_bench.rb
Rehearsal --------------------------------------------------------
naive map 4.370000 0.030000 4.400000 ( 4.458491)
naive map recursive 9.730000 0.090000 9.820000 ( 10.128538)
naive map via inject 7.550000 0.070000 7.620000 ( 7.766988)
---------------------------------------------- total: 21.840000sec
user system total real
naive map 4.360000 0.020000 4.380000 ( 4.413264)
naive map recursive 9.480000 0.050000 9.530000 ( 9.553978)
naive map via inject 7.420000 0.050000 7.470000 ( 7.509197)
..........................................................................
Do these numbers surprise you? As we expected, our inject based solution is
much faster than our recursive solution, but why is it so much slower than
our dumb brute force and ignorance approach?
The reason behind this is one of the key roadblocks that prevent us from
writing stateless code in Ruby. In order to solve this problem without
modifying any objects, we need to create a new object every single time an
element gets added to the array. As you may have guessed, objects are large in
Ruby, and constructing them is a slow proces. What's more, if we don't store any
of these intermediate values, we risk getting the garbage collector churning
frequently to kill off our discarded objects.
Avoiding side effects is different than avoiding mutable state entirely.
That's the key point to take away from what we just looked at here. In Ruby,
so long as it makes sense to do so, avoiding side effects is a good thing.
It reduces the possibility for unexpected bugs much in the same way that
avoiding the use of global variables does. However, avoiding the use of
mutable state definitely depends more on your individual situation.
We showed two examples that avoided the use of mutable state, both which might
look appealing to people who enjoy functional programming style. However,
we saw that their performance was abysmal, and because something like
`Enumerable#map` tends to be used in a tight loop, this is a bad time to
trade performance for aesthetic value.
However, in other situations, the trade-off may be tipped more in the other
direction. If the stateless (possibly recursive) code looks better than
other solutions, and performance is not a major concern, don't be afraid
to write your code in the more elegant way.
In general, remember the following things:
* The simple way to avoid side effects in Ruby when transforming one
object to another is to create a new object, and then populate it by
iterating over your original object performing the necessary state
transformations.
* You can write stateless code in Ruby by creating new objects every time
you perform an operation, e.g. `Array#+`.
* Recursive solutions may aid in writing simple stateless solutions, but
incur a major performance penalty in Ruby
* Creating too many objects can create performance problems as well,
so it is important to find the right balance, and to remember that
side effects can be avoided without making things fully stateless
We'll now move on from how you structure individual functions in your code
to how you can organize the larger chunks. So let's take a look at what we
can learn from modular organization, and how we can mix it in with our
object oriented code.
=== Modular Code Organization ===
In many functional languages, it is possible to group together your related
functions using a module. However, we typically think of something different
when we think of modules in Ruby:
..........................................................................
class A
include Enumerable
def initialize(arr)
@arr = arr
end
def each
@arr.each { |e| yield(e) }
end
end
>> A.new([1,2,3]).map { |x| x + 1 }
=> [2, 3, 4]
..........................................................................
Here, we've included the `Enumerable` module into our class as a mixin. This
enables shared implementation of functionality between classes, but is a
different concept than modular code organization in general.
What we really want is a collection of functions unified under a single
namespace. As it turns out, Ruby has that sort of thing too! Although
the `Math` module can be mixed into classes similar to the way we've used
`Enumerable` here, you can also use it on it own:
..........................................................................
>> Math.sin(Math::PI / 2)
=> 1.0
>> Math.sqrt(4)
=> 2.0
..........................................................................
So, how'd they do that? One way is to use `module_function`:
..........................................................................
module A
module_function
def foo
"This is foo"
end
def bar
"This is bar"
end
end
..........................................................................
We can now call these functions directly on the module, as you can see here:
..........................................................................
>> A.foo
=> "This is foo"
>> A.bar
=> "This is bar"
..........................................................................
You won't need anything more for most cases where you want to execute
functions on a module. However, this approach does come with some
limitations, because it does not allow you to use private functions:
..........................................................................
module A
module_function
def foo
"This is foo calling baz: #{baz}"
end
def bar
"This is bar"
end
private
def baz
"hi there"
end
end
..........................................................................
Though it seems like our code is fairly intuitive, we'll quickly run into
an error once we try to call `A.foo`:
..........................................................................
>> A.foo
NameError: undefined local variable or method `baz' for A:Module
from (irb):33:in `foo'
from (irb):46
from /Users/sandal/lib/ruby19_1/bin/irb:12:in `<main>'
..........................................................................
For some cases, not being able to access private methods might not be a big
deal, but for others, this could be a major issue. Luckily, if we think
laterally, there is an easy workaround.
Modules in Ruby, although they cannot be instantiated, are in essence
ordinary objects. Because of this, there is nothing stopping us from mixing
a module into itself.
..........................................................................
module A
extend self
def foo
"This is foo calling baz: #{baz}"
end
def bar
"This is bar"
end
private
def baz
"hi there"
end
end
..........................................................................
Once we do this, we get the same effect as `module_function` without the
limitations:
..........................................................................
>> A.foo
=> "This is foo calling baz: hi there"
..........................................................................
We aren't sacrificing encapsulation here, either. We will still get an
error if we try to call `A.baz` directly:
..........................................................................
>> A.baz
NoMethodError: private method `baz' called for A:Module
from (irb):65
from /Users/sandal/lib/ruby19_1/bin/irb:12:in `<main>'
..........................................................................
Using this trick of extending a module with itself provides us a structure
that isn't too different (at least on the surface) from the sort of modules
you might find in functional programming languages. But aside from odd cases
such as the `Math` module, you might wonder when this technique would be useful.
For the most part, classes work fine for encapsulating code in Ruby.
Traditional inheritance combined with the powerful mixin functionality of
modules covers most of the bases just fine. However, there are definitely
cases in which a concept isn't big enough for a class, but isn't small
enough to fit in a single function.
I ran into this issue recently in a Rails app I was working on. I was
implementing user authentication and needed to first test the database via
ActiveRecord, and fall back to LDAP when a user didn't have an account set
up in the application.
Without getting into too much detail, the basic structure for my
authentication routine looked like this:
..........................................................................
class User < ActiveRecord::Base
# other model code omitted
def self.authenticate(login, password)
if u = find_by_login(login)
u.authenticated?(password) ? u : nil
else
ldap_authenticate(login, password)
end
end
end
..........................................................................
LDAP authentication was to be implemented in a private class method, which
seemed like a good idea at first. However, as I continued to work on this,
I found myself writing a very huge function that represented over a page of
code. Since I knew there would be no way to keep this whole thing in my head
easily, I proceeded to break things into more helper methods to make things
clearer. This unfortunately didn't work as well as I had hoped.
By the end of this refactoring, I had racked up all sorts of strange routines
on `User`, with names such as `initialize_ldap_conn`, `retrieve_ldap_user`, etc.
A well factored object should do one thing and do it well, and my `User` model
seemed to know way too much about LDAP than it should have to. The
solution was to break this code off into a module, which was only a tiny
change to the `User.authenticate` method:
..........................................................................
def self.authenticate(login, password)
if u = find_by_login(login) # need to get the salt
u.authenticated?(password) ? u : nil
else
LDAP.authenticate(login, password)
end
end
..........................................................................
By substituting the private method call on the `User` model with a modular